Operating business22% entry signalFull studyWalk

AI Agent Infrastructure & Evaluation

Prepared 2026-09-08 · 2,096 words

The most attractive-looking and least enterable market in the portfolio. Growth you cannot capture is not an asset.

The industry — Computer systems design and related services (except video game design and development)

Base industry report for 541514 →
Establishments · CanadaA
41,081
with employees
Under 10 employeesA
88%
most common size: 1–4

Of 41,081 Canadian establishments with employees, 88% have fewer than ten — an industry of very small operators.

Entry signal — what decides who wins here

Structure decides
Structure decides One thing must be true Execution decides

The binding constraint is not executional. Being better than the incumbent does not, by itself, get you in — this one is cleared with capital, an asset, or a permission.

How it was read
Researched verdictUNVERIFIEDwalk — A full study: four structured dimensions, three kill criteria and a 30-day test behind the call.
How many new establishments are still tradingA
Professional, Scientific, and Technical Services, US · opened 2020
83.3%
1 year
64%
3 years
50.8%
5 years
34.3%
10 years
opened 2015

Measured, not forecast: the share of US establishments opening in one year that were still active later. It counts good operators and bad ones together, which is exactly why it is the honest answer to “what are the odds”. It is for the whole sector rather than this market, and the ten-year figure comes from an older cohort because no younger one has reached ten years.

This is not a probability of success, and it is not a verdict on you. No survival probability is published per market, and inventing one would be worse than saying so. What the bar reads is how much of the outcome sits inside an operator's control: green means the hurdles are ones a better operator clears, red means the binding constraint is capital, an asset or a permission rather than execution. Someone arriving with an advantage this screen did not assume can win a market shown in red.

The proposition being tested

Entering computer systems design and related services (except video game design and development) with Tracing, evaluation and regression-testing platform for LLM agents for AI engineering teams shipping agents to production.

No 30-day test. No 30-day test is proposed. Four conditions would have to be true to win and at least two are already false; running a test here would be theatre.
Market scaleinternationalunit: one developer, anywhere

Developer tools sell globally from day one with no geographic boundary whatsoever — which cuts both ways: the market is the whole world, and so is every competitor in it.

Boundary mechanismB None — distribution is a package registry
Angel-backed companies130
in the Canadian portfolio dataset
Province mixON 43, QC 34, AB 19, NB 9, NS 8, NL 6, BC 4, PE 3, US 2, SK 1, — 1

Sectors joined: Cybersecurity · SaaS · Technology · Dev Tools · B2B SaaS · AI

[UNVERIFIED] Sector-to-NAICS mapping is analyst judgment — see data/angel-sector-map.json. Counts are a per-record cross-reference and are not additive across records.

Screen score

6.45
Market size 5
Growth 10
Pain acuity 6
Incumbent vulnerability 3
Entry cost(inv) 8
Distribution access 4
Regulatory drag(inv) 9

Analyst judgment calibrated to the cited evidence, not measurement. Method

D

Demand landscape

Addressable market, competitor positions, and where buyer preference is shifting.

TAM — LLMOps market by 2028$4.8–4.9B (42% CAGR from $1.97B in 2024)
B

Two independent estimates — MarketsandMarkets and Bessemer — landing within $100M of each other. Unusually good corroboration for a young category, and the number is still small.

SAM — serviceable$300M
UNVERIFIED

A generous ~6% slice of the 2028 market after Datadog and the model providers take enterprise and bundled share — divided across 16+ funded competitors.

SOM — realistic capture$0–$96M

Even a strong 2% share of the total market is under $100M, against competitors who have already raised more than a new entrant ever will.

Demand indicators

Braintrust Series B, Feb 2026B$80M at $800M valuation
Langfuse Series BB$50M; tripling US sales and CS headcount; US ≈55% of revenue
Named funded competitors across three tiersB16+, before hyperscalers or open source
Capital into two direct competitors in ~12 monthsB$130M
Open-source free tierBLangfuse is open source — entry price for a comparable product is $0

Competitor positions

Observability tierno published share

Langfuse, LangSmith, Arize/Phoenix, Datadog LLM Observability, Helicone, Laminar, Latitude.

Evaluation & testing tierno published share

Braintrust, Patronus AI, Galileo, Giskard.

Governance & compliance tierno published share

Credo AI, Fiddler AI, Arthur AI, Holistic AI, Monitaur. Thinnest-funded tier — the one genuine adjacent opportunity.

Model providers (native tooling)no published share

Ship tracing and evaluation in-platform, free, integrated by default. The decisive competitor.

No reliable share data exists for a category this young. The competitor COUNT is the decision-relevant metric here, not share.

Shifting buyer preferences

  • Buyers increasingly accept bundled observability inside an existing APM contract rather than a separate tool.
  • Model providers' native eval and tracing tooling improves every release and costs nothing.
  • The genuine unmet need is cross-vendor portable evaluation that survives a model swap — a feature, not a company, and shippable by an incumbent in a quarter.
  • Buyer and builder are the same person: a market where the customer can plausibly rebuild your product in a weekend has a structural price ceiling.
R

Revenue model

Pricing that a real buyer would clear, the volume that follows, and what else the same customer will pay for.

Pricing

Open-source self-host$0

The competitive floor, set by a $50M-funded rival.

Team / cloud$6k–$25k
Enterprise$50k–$150k

Where the funded players compete hardest and a new entrant has no procurement credibility.

Average ticket — ACV$6k–$25k
UNVERIFIED

Structurally capped by a credible free tier from open source, free bundling from Datadog, and free native tooling from model providers. No amount of product quality changes this.

Volume projection

No projection is offered. Producing one would lend false precision to a market this study recommends walking away from.

Ancillary revenue

None material

Professional services in developer tooling are margin-dilutive and do not compound.

C

Cost structure

What it costs to stand this up and keep it running — and where the supply chain can end the business.

Fixed costs, annual

Engineering salaries$136k–$300k

Per head at the US software developer median, against teams funded at $80M and $50M.

Cloud and evaluation compute$25k–$150k

Running evals at customer scale is a genuine COGS line, unlike most SaaS.

SOC 2 / ISO 27001$30k–$60k

Table stakes for the enterprise tier.

Capital intensitymedium

Variable costs

Inference cost per evaluation run

LLM-as-judge means the product's own COGS scales with usage and is exposed to model pricing.

Paid acquisition

High search volume in a 16-competitor category means expensive keywords, not opportunity.

Supply chain

The model providers are simultaneously the supplier, the pricing authority, and the competitor. They set inference cost, they can change terms, and they ship the free alternative. There is no worse supply-chain position in this portfolio.

Labour — Canadian and US medians

RoleCA medianUS median
Software Developers

US employment 1,687,890.

$100,006$135,980
Data Scientists

US employment 262,440.

$95,992$120,230
Computer and Information Research Scientists

US employment 37,200.

—$140,300

The buyer is the same person as the builder, at the same wage. That symmetry caps price and raises cost simultaneously.

X

Execution & risk factors

Regulatory hurdles, whether anything defends the position once it works, and the macro trends acting on it.

Regulatory — low
Almost none — which is part of the problem, since regulatory drag is also a barrier to entry. The exception is the EU AI Act, which creates the governance-tier opening described in the verdict.
Defensibility — very low
No data moat, no switching cost, no network effect, and an open-source equivalent already funded at $50M.

Macro trends

Platform absorption from APM vendorsheadwind

Datadog bundles LLM observability into contracts that already exist. Observability niches have repeatedly ended this way.

Native tooling from model providersheadwind

Free, integrated, improving every release. Will not stop.

AI adoption growthtailwind

Real, and captured by incumbents rather than entrants.

EU AI Act compliance obligationstailwind

The only genuine opening — and it points at a different market (governance, closer to NAICS 5416), not a rescue of this one.

K

Kill criteria

The findings that should end this today. Written on the assumption that the reader is too invested to see them unaided.

KILL 1

$130M into two direct competitors in twelve months — $80M at an $800M valuation and $50M with a tripling US GTM team. Hiring, pricing patience and procurement credibility are all unmatched.

KILL 2

The category is being absorbed from both directions: Datadog bundles down from APM, model providers ship free native tooling up from the model.

KILL 3

Market size divided by competitor count: ~$4.8–4.9B by 2028 across 16+ funded players. Even 2% share is under $100M. The prize does not justify the fight.

#

Where the industry talks

The associations, forums and events where people in this trade actually talk shop — where to listen before entering, and where the first customers are found. Each link was opened on the date shown.

EventInternationalA
AI Engineer
ai.engineer

Conference series for software engineers building with AI; its page describes itself as the community and conference series for AI engineers.

Checked 2026-09-22
GroupInternationalA
MLOps Community
mlops.community

Slack community and meetup network; its home page invites you to join 70,000+ ML engineers. Where production eval and tracing practice gets argued out.

Checked 2026-09-22
PodcastInternationalA
Latent Space
latent.space

Newsletter and technical podcast on how labs build agents, models and infrastructure; published on Substack.

Checked 2026-09-22
ForumInternationalA
Hugging Face Forums
discuss.huggingface.co

Open Discourse forum for model, inference and evaluation questions.

Checked 2026-09-22
ForumInternationalA
LangChain Forum
forum.langchain.com

Discourse forum for LangChain and LangSmith users - the tracing and evaluation stack this record is about.

Checked 2026-09-22

Reddit returned no feed for r/LLMDevs or r/AI_Agents from this network, so no subreddit is listed here.

↔

Software serving this industry

Vertical software markets filed along the same branch of NAICS — who sells to these businesses, and who an entrant would have to displace.

Other records in this industry

§

Full study

The complete written report.

Market-Entry Study — AI Agent Infrastructure & Evaluation

NAICS 541514 · Computer systems design and related services (except video game design)

Verdict: WALK — this is the most attractive-looking and least enterable market in the portfolio. Prepared 2026-09-08 · Evidence tiers per ../_method/screening-model.md


The proposition being tested

Entering LLM/agent observability and evaluation tooling with a tracing, evaluation and regression-testing platform for AI engineering teams shipping agents to production.

This study exists to be honest about the market a technical operator will most want to enter, and to give three specific reasons not to.


1. MARKET SIZE

It is growing fast and it is small

Metric Value Tier
LLMOps market, 2024 $1.97B [B] MarketsandMarkets
Projected 2028 $4.9B (42% CAGR) [B]
Independent estimate — LLMOps incl. observability, eval, prompt mgmt, gateway/routing, cost optimisation, by 2028 $4.8B [B] Bessemer Venture Partners, Feb 2026

Two independent estimates landing within $100M of each other is unusually good corroboration for a category this young. Take ~$4.8–4.9B by 2028 as reliable.

The ratio that decides this

Now divide it. Named, funded players in the three tiers [B]:

Tier Players
Observability Langfuse, LangSmith, Arize AI / Phoenix, Datadog LLM Observability, Helicone, Laminar, Latitude
Evaluation & testing Braintrust, Patronus AI, Galileo, Giskard
Governance & compliance Credo AI, Fiddler AI, Arthur AI, Holistic AI, Monitaur

Sixteen named competitors before counting the hyperscalers, the model providers, or the open-source projects. A $4.8B market in 2028 divided across sixteen-plus funded teams, with Datadog taking the enterprise share it always takes in observability, leaves a median outcome well below the scale any of them raised against.

Compare to 238220: a US market of $3.1B growing 7.7% [B] where the leader holds 10.70% of installs [B]. Slower growth, comparable size, a fraction of the competitive density. Growth rate is the most seductive and least decision-relevant number in a market study.

Recent funding — the real barrier

Company Round Tier
Braintrust $80M Series B at $800M valuation, Feb 2026 [B]
Langfuse $50M Series B; Berlin HQ, opening SF + Singapore; US ≈55% of revenue; tripling US sales and CS headcount [B]

That last detail is the one to sit with. A competitor is tripling its go-to-market team in the exact geography a new entrant would target. This is not a market where a better product wins quietly.

Langfuse is also open source. The entry price for a comparable product is $0 plus a $50M-funded commercial arm.

Cost side, from this repo's own data [A, ../../occupation]

Occupation US employment US median CA median
Software developers 1,687,890 $135,980 $100,006
Data scientists 262,440 $120,230 $95,992
Computer & information research scientists 37,200 $140,300 —

Building competitively here means hiring against companies paying above these medians with $80M and $50M of fresh capital. Note also the buyer is the same person as the builder — a market where the customer can plausibly build your product in a weekend is a market with a structural price ceiling.

Demand signals

  • Market forecasts: STRONG and corroborated. Two independent sources agree.
  • Funding data: STRONG and negative. $130M into two direct competitors in roughly twelve months.
  • Search volume: NOT MEASURED. Run if pursuing anyway: LLM observability, agent evaluation, LLM eval platform, Langfuse alternative, Braintrust alternative. Expect high volume — and note that high volume in a category with sixteen funded players means expensive paid acquisition, not opportunity.
  • Reddit / HN: high discussion volume, but this is a builder population, not a buyer population. Enthusiasm here is a poor proxy for willingness to pay.
  • Amazon: NOT APPLICABLE.

Growing or shrinking: growing at ~42% CAGR [B], and consolidating simultaneously. Both at once. The second matters more.


2. THE CUSTOMER

What they want that nobody is giving them

The genuine unmet need — and it is genuine — is cross-vendor, portable evaluation that survives a model swap. Teams change models frequently and lose their evaluation history each time. Nobody solves this well, because every vendor has an incentive toward lock-in.

That is a real gap. It is also a feature, not a company, and the incumbents can ship it in a quarter.

What they pay for right now to solve it badly

Current spend Typical cost
Langfuse / LangSmith / Braintrust subscriptions $0 (OSS) → $50k+/yr enterprise
Datadog LLM Observability Bundled into existing APM spend — often effectively free at the margin
Internal eval harnesses 0.5–2 engineers; US median $135,980 [A]
Spreadsheets and manual review Free, ubiquitous, and the actual incumbent
Model providers' native tracing and eval tooling Free and improving

The last row is decisive. Every major model provider ships tracing and evaluation in-platform, at no marginal cost, integrated by default.

How much would they pay

[UNVERIFIED, and structurally capped.] With a credible free tier from open source, a free bundled option from the observability platform they already pay for, and free native tooling from the model provider, the price ceiling is set by competitors who have raised $130M and can price at zero indefinitely.

This is the definition of an unattractive pricing environment, and no amount of product quality changes it.


3. THE COMPETITION

Where they are slow, weak, or hated — steel-manning the entry

To be fair to the case for entering, real weaknesses exist:

  • Fragmentation. Teams commonly run one tool for tracing, another for eval, a third for governance. Consolidation is a genuine unmet need.
  • Open source complexity. Langfuse self-hosting has real operational cost.
  • Governance tier is thinnest. Credo AI, Fiddler, Arthur, Holistic AI and Monitaur are less capitalised than the observability tier, and EU AI Act obligations create a compliance-driven buyer with a deadline.
  • Eval quality is genuinely unsolved. LLM-as-judge is unreliable, and everybody knows it.

Why the gaps do not stay open

Every one of them is closeable by a funded incumbent in one to two quarters, and several are closeable by a model provider for free. The structural problem is not that competitors are strong — it is that the category is being absorbed from both directions at once: down from Datadog and the APM vendors, and up from the model providers. Observability niches have historically ended this way.

The single genuine exception is the governance tier, and it is not really this market — see the verdict.


4. ENTRY STRATEGY

Presented as required, then rejected. The best of these is still not good enough.

#1 — Vertical eval for a regulated domain. Cost: $80k–$200k. Odds: best of three. Evaluation and audit evidence for AI in healthcare, financial services or legal, where the buyer is a compliance officer with a regulatory deadline, not an engineer with a free alternative. Note what this actually is: a regulatory compliance business that happens to use eval technology. It escapes the pricing trap by leaving the market described in this study.

#2 — Open-source tool with a commercial tier. Cost: $50k + 12–18 months. Odds: low. The proven path in this category — and Langfuse already ran it, to $50M and a tripled US sales team. Following two years behind a well-funded incumbent on its own strategy is not a strategy.

#3 — Full observability + eval platform. Cost: $2M+. Odds: near zero. Direct competition with Braintrust at $800M and Datadog's distribution. Do not.

What would have to be true to win

  1. The market consolidates to 3–4 winners and a new entrant is one of them, from behind, without comparable capital.
  2. Model providers stop improving native eval tooling. They will not.
  3. Datadog fails to bundle LLM observability effectively. It already has.
  4. Buyers pay meaningfully for something with a credible free tier.

Four required conditions, at least two of which are already false. That is a walk, and no 30-day test changes it — which is why none is proposed. Running one here would be theatre.


5. KILL CRITERIA

1. $130M into two direct competitors in twelve months. Braintrust at $80M / $800M valuation, Langfuse at $50M and tripling US GTM [B]. A new entrant cannot match hiring, cannot match pricing patience, and cannot match the enterprise credibility a $800M valuation buys in a procurement conversation.

2. The category is being absorbed from both directions. Datadog bundles LLM observability into APM contracts that already exist; model providers ship native tracing and eval free. Independent observability categories have repeatedly been squeezed out by exactly this pincer. There is no reason this one is the exception.

3. Market size divided by competitor count. ~$4.8–4.9B by 2028 [B] across sixteen-plus named funded players plus hyperscalers plus open source. Even a strong outcome — 2% share — is under $100M in a market where the leaders have already raised more than a new entrant will ever put in. The prize does not justify the fight.

The honest bias check, and the reason this study is in the portfolio: this is the market a technical operator will most want to enter. The tooling is familiar, the problems are legible, the customers are people like you, the vocabulary is comfortable, and building the first version would be genuinely enjoyable. Every one of those is a reason to be more suspicious, not less. Familiarity with a market is not an advantage in it — sixteen other teams have the same familiarity and $130M more capital. Meanwhile the two markets this portfolio actually recommends — interconnection queues and HVAC service agreements — are boring, and boring is what is left uncontested.


THE CALL: WALK

Walk. Not "wait" — there is no trigger that would reopen this. The structural conditions worsen with time as incumbents consolidate and native tooling improves.

The evidence: two corroborated forecasts putting the 2028 market at ~$4.8–4.9B [B]; sixteen-plus named funded competitors across three tiers [B]; $130M raised by two direct competitors in the last twelve months, one at an $800M valuation and one tripling its US go-to-market team [B]; a credible open-source free tier; free bundled competition from Datadog; and free native tooling from the model providers. Four conditions would have to be true to win, and two are already false.

The one adjacent thing worth doing

If AI evaluation must be pursued, pursue AI governance and compliance evidence for regulated industries — the thinnest-funded tier, with a compliance buyer facing a regulatory deadline rather than an engineer with a free alternative. That is a different market (closer to NAICS 5416, management and technical consulting) and deserves its own study before any commitment. It is not a rescue of this one.

Do not revisit this market. Reallocate the attention to 221121.


STRUCTURED ANALYSIS

Four dimensions of this study — demand landscape, revenue model, cost structure, and execution & risk factors — are held as structured data in profile.json in this folder rather than repeated as prose here, so there is exactly one source of truth for every figure.

Dimension What it holds
demand TAM / SAM / SOM with evidence tiers, demand indicators, competitor positions and published shares where they exist, shifting buyer preferences
revenue Pricing tiers, average ticket, five-year volume and revenue projection, ancillary revenue streams
cost Fixed and variable operating costs, capital intensity, supply-chain dependency, and labour medians drawn from the Occupation Atlas
risk Regulatory level, defensibility, and macro trends tagged tailwind / headwind / mixed

The Market Research app renders all four as panels above this report — run npm run dev from markets/, or open /reports/<naics>.


Sources