AI Agent Infrastructure & Evaluation
The most attractive-looking and least enterable market in the portfolio. Growth you cannot capture is not an asset.
The industry — Computer systems design and related services (except video game design and development)
Base industry report for 541514 →- Establishments · CanadaA
- 41,081
- Under 10 employeesA
- 88%
Of 41,081 Canadian establishments with employees, 88% have fewer than ten — an industry of very small operators.
Entry signal — what decides who wins here
Structure decidesThe binding constraint is not executional. Being better than the incumbent does not, by itself, get you in — this one is cleared with capital, an asset, or a permission.
Measured, not forecast: the share of US establishments opening in one year that were still active later. It counts good operators and bad ones together, which is exactly why it is the honest answer to “what are the odds”. It is for the whole sector rather than this market, and the ten-year figure comes from an older cohort because no younger one has reached ten years.
This is not a probability of success, and it is not a verdict on you. No survival probability is published per market, and inventing one would be worse than saying so. What the bar reads is how much of the outcome sits inside an operator's control: green means the hurdles are ones a better operator clears, red means the binding constraint is capital, an asset or a permission rather than execution. Someone arriving with an advantage this screen did not assume can win a market shown in red.
The proposition being tested
Entering computer systems design and related services (except video game design and development) with Tracing, evaluation and regression-testing platform for LLM agents for AI engineering teams shipping agents to production.
Developer tools sell globally from day one with no geographic boundary whatsoever — which cuts both ways: the market is the whole world, and so is every competitor in it.
Sectors joined: Cybersecurity · SaaS · Technology · Dev Tools · B2B SaaS · AI
[UNVERIFIED] Sector-to-NAICS mapping is analyst judgment — see data/angel-sector-map.json. Counts are a per-record cross-reference and are not additive across records.
Screen score
6.45Analyst judgment calibrated to the cited evidence, not measurement. Method
Demand landscape
Addressable market, competitor positions, and where buyer preference is shifting.
Two independent estimates — MarketsandMarkets and Bessemer — landing within $100M of each other. Unusually good corroboration for a young category, and the number is still small.
A generous ~6% slice of the 2028 market after Datadog and the model providers take enterprise and bundled share — divided across 16+ funded competitors.
Even a strong 2% share of the total market is under $100M, against competitors who have already raised more than a new entrant ever will.
Demand indicators
Competitor positions
Langfuse, LangSmith, Arize/Phoenix, Datadog LLM Observability, Helicone, Laminar, Latitude.
Braintrust, Patronus AI, Galileo, Giskard.
Credo AI, Fiddler AI, Arthur AI, Holistic AI, Monitaur. Thinnest-funded tier — the one genuine adjacent opportunity.
Ship tracing and evaluation in-platform, free, integrated by default. The decisive competitor.
No reliable share data exists for a category this young. The competitor COUNT is the decision-relevant metric here, not share.
Shifting buyer preferences
- Buyers increasingly accept bundled observability inside an existing APM contract rather than a separate tool.
- Model providers' native eval and tracing tooling improves every release and costs nothing.
- The genuine unmet need is cross-vendor portable evaluation that survives a model swap — a feature, not a company, and shippable by an incumbent in a quarter.
- Buyer and builder are the same person: a market where the customer can plausibly rebuild your product in a weekend has a structural price ceiling.
Revenue model
Pricing that a real buyer would clear, the volume that follows, and what else the same customer will pay for.
Pricing
The competitive floor, set by a $50M-funded rival.
Where the funded players compete hardest and a new entrant has no procurement credibility.
Structurally capped by a credible free tier from open source, free bundling from Datadog, and free native tooling from model providers. No amount of product quality changes this.
Volume projection
No projection is offered. Producing one would lend false precision to a market this study recommends walking away from.
Ancillary revenue
Professional services in developer tooling are margin-dilutive and do not compound.
Cost structure
What it costs to stand this up and keep it running — and where the supply chain can end the business.
Fixed costs, annual
Per head at the US software developer median, against teams funded at $80M and $50M.
Running evals at customer scale is a genuine COGS line, unlike most SaaS.
Table stakes for the enterprise tier.
Variable costs
LLM-as-judge means the product's own COGS scales with usage and is exposed to model pricing.
High search volume in a 16-competitor category means expensive keywords, not opportunity.
Supply chain
The model providers are simultaneously the supplier, the pricing authority, and the competitor. They set inference cost, they can change terms, and they ship the free alternative. There is no worse supply-chain position in this portfolio.
Labour — Canadian and US medians
| Role | CA median | US median |
|---|---|---|
| Software Developers US employment 1,687,890. | $100,006 | $135,980 |
| Data Scientists US employment 262,440. | $95,992 | $120,230 |
| Computer and Information Research Scientists US employment 37,200. | — | $140,300 |
The buyer is the same person as the builder, at the same wage. That symmetry caps price and raises cost simultaneously.
Execution & risk factors
Regulatory hurdles, whether anything defends the position once it works, and the macro trends acting on it.
Macro trends
Datadog bundles LLM observability into contracts that already exist. Observability niches have repeatedly ended this way.
Free, integrated, improving every release. Will not stop.
Real, and captured by incumbents rather than entrants.
The only genuine opening — and it points at a different market (governance, closer to NAICS 5416), not a rescue of this one.
Kill criteria
The findings that should end this today. Written on the assumption that the reader is too invested to see them unaided.
$130M into two direct competitors in twelve months — $80M at an $800M valuation and $50M with a tripling US GTM team. Hiring, pricing patience and procurement credibility are all unmatched.
The category is being absorbed from both directions: Datadog bundles down from APM, model providers ship free native tooling up from the model.
Market size divided by competitor count: ~$4.8–4.9B by 2028 across 16+ funded players. Even 2% share is under $100M. The prize does not justify the fight.
Where the industry talks
The associations, forums and events where people in this trade actually talk shop — where to listen before entering, and where the first customers are found. Each link was opened on the date shown.
Conference series for software engineers building with AI; its page describes itself as the community and conference series for AI engineers.
Slack community and meetup network; its home page invites you to join 70,000+ ML engineers. Where production eval and tracing practice gets argued out.
Newsletter and technical podcast on how labs build agents, models and infrastructure; published on Substack.
Open Discourse forum for model, inference and evaluation questions.
Discourse forum for LangChain and LangSmith users - the tracing and evaluation stack this record is about.
Reddit returned no feed for r/LLMDevs or r/AI_Agents from this network, so no subreddit is listed here.
Software serving this industry
Vertical software markets filed along the same branch of NAICS — who sells to these businesses, and who an entrant would have to displace.
Other records in this industry
Full study
The complete written report.
Market-Entry Study — AI Agent Infrastructure & Evaluation
NAICS 541514 · Computer systems design and related services (except video game design)
Verdict: WALK — this is the most attractive-looking and least enterable market in the portfolio.
Prepared 2026-09-08 · Evidence tiers per ../_method/screening-model.md
The proposition being tested
Entering LLM/agent observability and evaluation tooling with a tracing, evaluation and regression-testing platform for AI engineering teams shipping agents to production.
This study exists to be honest about the market a technical operator will most want to enter, and to give three specific reasons not to.
1. MARKET SIZE
It is growing fast and it is small
| Metric | Value | Tier |
|---|---|---|
| LLMOps market, 2024 | $1.97B | [B] MarketsandMarkets |
| Projected 2028 | $4.9B (42% CAGR) | [B] |
| Independent estimate — LLMOps incl. observability, eval, prompt mgmt, gateway/routing, cost optimisation, by 2028 | $4.8B | [B] Bessemer Venture Partners, Feb 2026 |
Two independent estimates landing within $100M of each other is unusually good corroboration for a category this young. Take ~$4.8–4.9B by 2028 as reliable.
The ratio that decides this
Now divide it. Named, funded players in the three tiers [B]:
| Tier | Players |
|---|---|
| Observability | Langfuse, LangSmith, Arize AI / Phoenix, Datadog LLM Observability, Helicone, Laminar, Latitude |
| Evaluation & testing | Braintrust, Patronus AI, Galileo, Giskard |
| Governance & compliance | Credo AI, Fiddler AI, Arthur AI, Holistic AI, Monitaur |
Sixteen named competitors before counting the hyperscalers, the model providers, or the open-source projects. A $4.8B market in 2028 divided across sixteen-plus funded teams, with Datadog taking the enterprise share it always takes in observability, leaves a median outcome well below the scale any of them raised against.
Compare to 238220: a US market of $3.1B growing 7.7% [B] where the leader holds 10.70% of installs [B]. Slower growth, comparable size, a fraction of the competitive density. Growth rate is the most seductive and least decision-relevant number in a market study.
Recent funding — the real barrier
| Company | Round | Tier |
|---|---|---|
| Braintrust | $80M Series B at $800M valuation, Feb 2026 | [B] |
| Langfuse | $50M Series B; Berlin HQ, opening SF + Singapore; US ≈55% of revenue; tripling US sales and CS headcount | [B] |
That last detail is the one to sit with. A competitor is tripling its go-to-market team in the exact geography a new entrant would target. This is not a market where a better product wins quietly.
Langfuse is also open source. The entry price for a comparable product is $0 plus a $50M-funded commercial arm.
Cost side, from this repo's own data [A, ../../occupation]
| Occupation | US employment | US median | CA median |
|---|---|---|---|
| Software developers | 1,687,890 | $135,980 | $100,006 |
| Data scientists | 262,440 | $120,230 | $95,992 |
| Computer & information research scientists | 37,200 | $140,300 | — |
Building competitively here means hiring against companies paying above these medians with $80M and $50M of fresh capital. Note also the buyer is the same person as the builder — a market where the customer can plausibly build your product in a weekend is a market with a structural price ceiling.
Demand signals
- Market forecasts: STRONG and corroborated. Two independent sources agree.
- Funding data: STRONG and negative. $130M into two direct competitors in roughly twelve months.
- Search volume: NOT MEASURED. Run if pursuing anyway:
LLM observability,agent evaluation,LLM eval platform,Langfuse alternative,Braintrust alternative. Expect high volume — and note that high volume in a category with sixteen funded players means expensive paid acquisition, not opportunity. - Reddit / HN: high discussion volume, but this is a builder population, not a buyer population. Enthusiasm here is a poor proxy for willingness to pay.
- Amazon: NOT APPLICABLE.
Growing or shrinking: growing at ~42% CAGR [B], and consolidating simultaneously. Both at once. The second matters more.
2. THE CUSTOMER
What they want that nobody is giving them
The genuine unmet need — and it is genuine — is cross-vendor, portable evaluation that survives a model swap. Teams change models frequently and lose their evaluation history each time. Nobody solves this well, because every vendor has an incentive toward lock-in.
That is a real gap. It is also a feature, not a company, and the incumbents can ship it in a quarter.
What they pay for right now to solve it badly
| Current spend | Typical cost |
|---|---|
| Langfuse / LangSmith / Braintrust subscriptions | $0 (OSS) → $50k+/yr enterprise |
| Datadog LLM Observability | Bundled into existing APM spend — often effectively free at the margin |
| Internal eval harnesses | 0.5–2 engineers; US median $135,980 [A] |
| Spreadsheets and manual review | Free, ubiquitous, and the actual incumbent |
| Model providers' native tracing and eval tooling | Free and improving |
The last row is decisive. Every major model provider ships tracing and evaluation in-platform, at no marginal cost, integrated by default.
How much would they pay
[UNVERIFIED, and structurally capped.] With a credible free tier from open source, a free bundled option from the observability platform they already pay for, and free native tooling from the model provider, the price ceiling is set by competitors who have raised $130M and can price at zero indefinitely.
This is the definition of an unattractive pricing environment, and no amount of product quality changes it.
3. THE COMPETITION
Where they are slow, weak, or hated — steel-manning the entry
To be fair to the case for entering, real weaknesses exist:
- Fragmentation. Teams commonly run one tool for tracing, another for eval, a third for governance. Consolidation is a genuine unmet need.
- Open source complexity. Langfuse self-hosting has real operational cost.
- Governance tier is thinnest. Credo AI, Fiddler, Arthur, Holistic AI and Monitaur are less capitalised than the observability tier, and EU AI Act obligations create a compliance-driven buyer with a deadline.
- Eval quality is genuinely unsolved. LLM-as-judge is unreliable, and everybody knows it.
Why the gaps do not stay open
Every one of them is closeable by a funded incumbent in one to two quarters, and several are closeable by a model provider for free. The structural problem is not that competitors are strong — it is that the category is being absorbed from both directions at once: down from Datadog and the APM vendors, and up from the model providers. Observability niches have historically ended this way.
The single genuine exception is the governance tier, and it is not really this market — see the verdict.
4. ENTRY STRATEGY
Presented as required, then rejected. The best of these is still not good enough.
#1 — Vertical eval for a regulated domain. Cost: $80k–$200k. Odds: best of three. Evaluation and audit evidence for AI in healthcare, financial services or legal, where the buyer is a compliance officer with a regulatory deadline, not an engineer with a free alternative. Note what this actually is: a regulatory compliance business that happens to use eval technology. It escapes the pricing trap by leaving the market described in this study.
#2 — Open-source tool with a commercial tier. Cost: $50k + 12–18 months. Odds: low. The proven path in this category — and Langfuse already ran it, to $50M and a tripled US sales team. Following two years behind a well-funded incumbent on its own strategy is not a strategy.
#3 — Full observability + eval platform. Cost: $2M+. Odds: near zero. Direct competition with Braintrust at $800M and Datadog's distribution. Do not.
What would have to be true to win
- The market consolidates to 3–4 winners and a new entrant is one of them, from behind, without comparable capital.
- Model providers stop improving native eval tooling. They will not.
- Datadog fails to bundle LLM observability effectively. It already has.
- Buyers pay meaningfully for something with a credible free tier.
Four required conditions, at least two of which are already false. That is a walk, and no 30-day test changes it — which is why none is proposed. Running one here would be theatre.
5. KILL CRITERIA
1. $130M into two direct competitors in twelve months. Braintrust at $80M / $800M valuation, Langfuse at $50M and tripling US GTM [B]. A new entrant cannot match hiring, cannot match pricing patience, and cannot match the enterprise credibility a $800M valuation buys in a procurement conversation.
2. The category is being absorbed from both directions. Datadog bundles LLM observability into APM contracts that already exist; model providers ship native tracing and eval free. Independent observability categories have repeatedly been squeezed out by exactly this pincer. There is no reason this one is the exception.
3. Market size divided by competitor count. ~$4.8–4.9B by 2028 [B] across sixteen-plus named funded players plus hyperscalers plus open source. Even a strong outcome — 2% share — is under $100M in a market where the leaders have already raised more than a new entrant will ever put in. The prize does not justify the fight.
The honest bias check, and the reason this study is in the portfolio: this is the market a technical operator will most want to enter. The tooling is familiar, the problems are legible, the customers are people like you, the vocabulary is comfortable, and building the first version would be genuinely enjoyable. Every one of those is a reason to be more suspicious, not less. Familiarity with a market is not an advantage in it — sixteen other teams have the same familiarity and $130M more capital. Meanwhile the two markets this portfolio actually recommends — interconnection queues and HVAC service agreements — are boring, and boring is what is left uncontested.
THE CALL: WALK
Walk. Not "wait" — there is no trigger that would reopen this. The structural conditions worsen with time as incumbents consolidate and native tooling improves.
The evidence: two corroborated forecasts putting the 2028 market at ~$4.8–4.9B [B]; sixteen-plus named funded competitors across three tiers [B]; $130M raised by two direct competitors in the last twelve months, one at an $800M valuation and one tripling its US go-to-market team [B]; a credible open-source free tier; free bundled competition from Datadog; and free native tooling from the model providers. Four conditions would have to be true to win, and two are already false.
The one adjacent thing worth doing
If AI evaluation must be pursued, pursue AI governance and compliance evidence for regulated industries — the thinnest-funded tier, with a compliance buyer facing a regulatory deadline rather than an engineer with a free alternative. That is a different market (closer to NAICS 5416, management and technical consulting) and deserves its own study before any commitment. It is not a rescue of this one.
Do not revisit this market. Reallocate the attention to 221121.
STRUCTURED ANALYSIS
Four dimensions of this study — demand landscape, revenue model, cost
structure, and execution & risk factors — are held as structured data in
profile.json in this folder rather than repeated as prose here,
so there is exactly one source of truth for every figure.
| Dimension | What it holds |
|---|---|
demand |
TAM / SAM / SOM with evidence tiers, demand indicators, competitor positions and published shares where they exist, shifting buyer preferences |
revenue |
Pricing tiers, average ticket, five-year volume and revenue projection, ancillary revenue streams |
cost |
Fixed and variable operating costs, capital intensity, supply-chain dependency, and labour medians drawn from the Occupation Atlas |
risk |
Regulatory level, defensibility, and macro trends tagged tailwind / headwind / mixed |
The Market Research app renders all four as panels above this report — run
npm run dev from markets/, or open /reports/<naics>.
Sources
- MarkTechPost — Top LLM Observability and Evaluation Platforms in 2026
- Deepak Gupta — AI Agent Observability & Governance: 2026 Market Reality Check
- Callsphere — AI Agent Observability Platform Langfuse Raises $50M Series B
- Latitude — AI agent observability tools compared (2026)
- Laminar — Top 6 Agent Observability Platforms (2026)
- Kanerika — LLMOps Observability: LangSmith vs Arize vs Langfuse vs W&B
- Wage data:
../../occupation/data/build/site-data.json(BLS OEWS / ESDC Job Bank 2025)