# What CFOs, CIOs and CTOs Weigh When Choosing an LLM

Enterprises buy LLMs on outcomes, not leaderboards or tokens: 85% treat benchmarks as background or let their own workload test decide, 60% judge cost by ROI, yet 53% have no cost-per-outcome metric. n=102.

**Analyst:** Anubha Garg

*Published by G2 Research · September 29, 2026 · Last updated October 5, 2026*

> Enterprises buy large language models (LLMs) for outcomes, yet 53% can't measure them. Based on G2 first-party research from 102 in-depth interviews with US enterprise leaders who evaluate, select or fund LLMs.

**Executive answer:** The CFOs, CIOs and CTOs who sign LLM contracts use neither leaderboard scores nor price per million tokens. They run the model on their own data, put it through a security review and judge the bill by the hours it saves. Of 93 leaders, 85% treated public benchmarks as background or let a test on their own workloads decide. 60% of 98 compare models on downstream return on investment (ROI), yet 53% of 96 have no formal cost-per-outcome metric. Security ends deals (59% of 75), 40% of 97 have been surprised by a regression or silent model change, 75% of 101 exclude or never tested non-US models, and 80% of 84 route work between models by hand.

**On this page:** [Benchmarks](#do-enterprise-buyers-actually-use-llm-benchmarks-to-choose-a-model) · [Cost per outcome](#is-token-cost-the-ceiling-on-enterprise-llm-scale) · [Production reliability](#what-does-working-well-in-production-mean-to-llm-buyers) · [Governance](#which-governance-and-admin-controls-decide-whether-an-llm-gets-deployed) · [Non-US models](#are-non-us-llms-such-as-deepseek-and-qwen-getting-into-the-enterprise) · [Multi-model routing](#how-do-enterprises-route-work-across-llm-providers-and-what-triggers-a-switch) · [Recommendations](#what-should-llm-providers-and-vendors-do) · [FAQ](#faq) · [Methodology](#methodology)

**Key terms.** *Public benchmarks* are leaderboards that rank LLMs on general tasks, not on any one company's data. *Token economics* is paying per million tokens; *cost per outcome* is the price of one finished task.

## Key findings

| Finding | Share | Base |
|---|---|---|
| Treated benchmarks as background or let their own workload test decide | 85% | 93 |
| Compare models on downstream ROI rather than price per token | 60% | 98 |
| Have no formal cost-per-outcome metric | 53% | 96 |
| Named security, privacy or data exposure as the near-dealbreaker | 59% | 75 |
| Last production surprise was a regression, hallucination or silent model change | 40% | 97 |
| Exclude or have never tested non-US models | 75% | 101 |
| Run multi-provider environments | 63% | 101 |
| Route work between models by hand | 80% | 84 |

---

## What do enterprise buyers actually weigh when they choose an LLM?

Buyers weigh four things: whether the model passes on their own data, whether it clears the security review, what one outcome costs and whether it behaves the same next month. Of the 93 leaders who described where public benchmarks sat in their last LLM decision, 85% either treated them as background or let a test on their own workloads decide.

Once a model passes on the buyer's data, the test is consistency and the gate is security. Half of buyers define a working model as one that completes tasks reliably every time, and 40% have already been surprised by a regression, a hallucination or a silent model change. Access, permissions and audit decide deployment, and security and legal teams hold the veto. Non-US models rarely reach the test at all, while multi-model estates are already the norm and are run by hand.

**G2 perspective:** "The teams scaling successfully today aren't choosing between AI adoption and token efficiency. They're embracing both while measuring cost per outcome. When 53% of enterprise leaders lack that metric, they manage inputs rather than outcomes. That gap explains why costs are already hitting the ceiling for one in four, and why those with clarity on the unit of work are moving forward." — Godard Abel, CEO, G2

**TL;DR:** Enterprises choose LLMs on fit, security, cost per outcome and predictability, not on leaderboard rank or token price.

### Six themes shaping the market

| Theme | Finding |
|---|---|
| Benchmarks | Benchmarks screen, workloads decide: the leaderboard earns a first look; the pilot on the buyer's own data makes the decision. |
| Cost per outcome | Tokens are the vocabulary, outcomes are the arithmetic: buyers trade tokens for hours of people's time; one in four has already hit a cost ceiling. |
| Reliability in production | Consistency is the production test: the failures buyers remember are the silent ones. |
| Security and governance | Security holds the veto: access, permissions and audit are the daily work. |
| Non-US models | The door is shut on security grounds, not performance: residency, IP and procurement rules exclude non-US models before any benchmark is read. |
| Multi-model routing | Portfolios are real, routers are not: most enterprises run several labs and route between them by hand. |

---

## Do enterprise buyers actually use LLM benchmarks to choose a model?

Rarely as the deciding factor. Of the 93 leaders who described where public benchmarks and leaderboards sat in their last model decision, 44% called them background context or ignored them outright, and 41% said testing on their own workloads was the decisive evidence. Together that is 85% for whom the leaderboard did not decide.

| Role of benchmarks in the last LLM decision | Respondents | Share |
|---|---|---|
| Background context or ignored | 41 | 44% |
| Real-workload testing was the decisive evidence | 38 | 41% |
| Worried vendor claims can fail in production | 14 | 15% |

*Base: 93 respondents.*

### What replaces the leaderboard?

Task capability and accuracy on the buyer's own use case came first among the criteria that carried the last decision (42% of 96). Security and integration (29%) and cost, speed and hands-on testing (29%) shared second place. Benchmarks were not on the list. A revenue operations VP in hospitality said: "I actually don't look at public benchmarks or leaderboards. At all."

### Have vendor claims failed on real work?

Yes, for 18% of the 96 leaders asked, who described a leaderboard result or vendor claim that did not survive contact with their workload. A footwear company found Copilot's results on real workloads "very different from what we expected." The other 77% had no such story, often because internal testing on live data comes first.

### What evidence do buyers want instead of benchmarks?

Real-workload pilots and direct testing on their own data, named by 42% of 92 leaders. Only 2% asked for better public evidence.

### Who evaluates the model?

Of 100 leaders, 61% are hands-on in evaluating and selecting models; the rest sit in a cross-functional governance role or own the budget. Decision rights are shared for 45%, so evidence has to convince both the engineer who ran the test and the finance or security owner who did not.

### What almost stops an LLM purchase?

Security. Of the 75 leaders who named a near-dealbreaker in their most recent model selection, 59% pointed to security, privacy or data exposure.

| Near-dealbreaker | Respondents | Share |
|---|---|---|
| Security, privacy and data-exposure risk | 44 | 59% |
| Failure on real-world task quality or functionality | 21 | 28% |
| Unfavorable pricing, token economics or redundant spend | 10 | 13% |

*Base: 75 respondents.*

> "We definitely found that some of the vendors cost for performance ratios were significantly more expensive than what was advertised once we got into the actual work flow and requirements. It was not necessarily misleading. I think it was overly optimistic." — CIO and CTO, healthcare technology

**G2 perspective:** "You might have a model that works well for most things, then find another that handles one particular task better or gets the same result for less. You don't need to rethink your whole setup. Test it on that task, see if the improvement holds up, and move some of the work over. I think that's a pretty sensible way to buy in a market that keeps changing." — Alex David, GM of AI Solutions, G2

**Strategic implication:** Treat the public leaderboard as the price of admission and the buyer's pilot as the sale. LLM providers should publish workload-level evidence for the use cases buyers name (customer support, coding, document processing and analytics) with cost and failure modes attached, and lead with the security review.

**TL;DR:** Benchmarks build the shortlist; a pilot on the buyer's own data makes the decision, and security can end it.

---

## Is token cost the ceiling on enterprise LLM scale?

For one buyer in four it already is. Of the 91 leaders who discussed cost per outcome, 26% described cost-driven limits, redesigns or model downgrades, and half of all buyers have no threshold at which the economics would tell them to scale or stop. Of 98 asked how they compare models on cost, 60% look at downstream ROI and labor savings rather than price per million tokens.

### Have enterprises adopted a cost-per-outcome unit?

Most have not: 53% of 96 leaders have no formal cost-per-outcome metric. Only 29% can price a finished task.

| Cost-per-outcome unit | Respondents | Share |
|---|---|---|
| No formal cost-per-outcome metric | 51 | 53% |
| Case, ticket, document or task-level measure | 28 | 29% |
| Time savings, labor avoidance and aggregate productivity | 17 | 18% |

*Base: 96 respondents.*

Where the unit exists, the arithmetic is short. A pharmaceutical governance lead weighs an automated report that "saves us four hours" and "costs us a thousand dollars" against a specialist whose wage rate is $400 an hour.

> "One of our largest uses is with Fin and Intercom. And they charge a fee of $1 per ticket that's touched. And when we look at the loaded labor rate of our human agents, it's a no brainer for us. $1 versus at least $50 for a human interaction when you consider everything." — Senior Director of AI and Automations, SaaS

### How is cost already limiting scale?

Through caps, switch-offs and closed projects. An enterprise SaaS team capped total annual spend after an "unexpected ramp up in cost"; a construction IT director "had to turn off different abilities based on token usage"; a technology CEO "closed down several projects because they were running above cost." The bill lands with finance, the CFO or the CIO for three-quarters of those who named an owner.

### Do buyers think in tokens?

They talk in tokens but reason in outcomes. Sixty-eight of the 102 leaders used the word "token" in their interview, yet only one reasoned from the advertised price per million tokens. The cheapest token price is not the best deal if the model needs a second pass or a human correction.

### Who measures cost per outcome?

Budget owners more than evaluators. Of the 18 leaders who own the budget or approve spend, 50% measure cost per case, ticket, document or task; of the 61 hands-on evaluators, 25% do. The budget-owner group is under 30 and directional.

| Decision role | Task-level measure | Time savings / productivity | No formal metric |
|---|---|---|---|
| Budget ownership or spend approval | 9 | 3 | 6 |
| Cross-functional governance and recommendation | 4 | 5 | 10 |
| Hands-on model evaluation and selection | 15 | 9 | 34 |

*Respondent counts. Segments under 30 are directional.*

### What happens without a unit?

Buyers police tokens instead of outcomes. Of the 28 leaders with a task-level unit, 54% have a rule for when to scale or stop; of the 51 without one, 41% do. 37% of those with no metric describe managing token usage, against 7% of those who measure per outcome. Both groups are under 30 and directional.

### What do buyers do when cost bites?

They downgrade the model or push for subscription pricing. Cost drove only 15% of the last model changes leaders described.

**Strategic implication:** Price the outcome, not the token, and publish the ceiling. Labs and platforms that publish a cost per resolved outcome for named workloads, with a forecastable ceiling, speak the unit finance uses.

**TL;DR:** Buyers judge LLM spend in hours saved, but most cannot measure cost per outcome, and one in four has already hit a cost ceiling.

---

## What does "working well in production" mean to LLM buyers?

Consistency. Of the 98 leaders asked what a model working well in production means to them, 52% described reliable, accurate and consistent task completion. A third defined it as measurable workflow efficiency and business outcomes, and the rest as adoption, satisfaction and low escalation rates.

### What was the last bad surprise in production?

For 40% of 97 leaders it was an output regression, a hallucination or a silent model change; three in four have been burned in production.

| Last production surprise | Respondents | Share |
|---|---|---|
| Output regressions, hallucinations or silent model changes | 39 | 40% |
| Outages, integration failures and unexpected cost spikes | 33 | 34% |
| No material production incident reported | 25 | 26% |

*Base: 97 respondents.*

> "I would say going back into a workflow, or agent model where I had refined and worked with it. And after a change or update, noticed that the output was slightly different or altered and need to go back and reevaluate to hone it in and get the expected result that I previously had got working." — Program Management, technology, software and hardware

### Do provider updates break production systems?

Sometimes, and few buyers are protected. Of the 78 asked directly about provider updates, 21% said a silent update or version change had caused a regression, and only 8% described testing, monitoring, rollback or fallback built for it. Of the 70 leaders with a vendor gripe, 34% told a story about silent updates, outages and poor change communication.

### Is consistency measured or judged by feel?

Mostly judged by feel. Of 67 leaders, 40% call consistency a core production requirement without naming a method, 37% assess it through user feedback, and 22% measure it with business outcomes, accuracy and dashboards.

| Approach to consistency | Respondents | Share |
|---|---|---|
| Consistency is a core production requirement (no method named) | 27 | 40% |
| Assessed qualitatively through user feedback and absence of regressions | 25 | 37% |
| Measured through business outcomes, accuracy and dashboards | 15 | 22% |

*Base: 67 respondents.*

**Strategic implication:** Make model changes visible and reversible. Version pinning, dated changelogs, regression suites on the buyer's own prompts and a fallback path are product features that decide renewals.

**TL;DR:** A working LLM is one that behaves the same next month, and silent model changes are the failure buyers remember.

---

## Which governance and admin controls decide whether an LLM gets deployed?

Access, permissions and sensitive-data governance. Asked what it takes to run a model day to day, setting quality and cost aside, 43% of 98 leaders put access and sensitive-data governance first, 35% named integration with enterprise systems, and 22% monitoring, adoption and support.

| Day-to-day operational work | Respondents | Share |
|---|---|---|
| Access, permissions and sensitive-data governance | 42 | 43% |
| Integration with enterprise systems and workflows | 34 | 35% |
| Monitoring, adoption and ongoing operational support | 22 | 22% |

*Base: 98 respondents.*

### Who holds veto power over an LLM decision?

Security, privacy, legal or compliance teams for 44% of the 77 leaders who named a veto holder, and a cross-functional governance group for 43%.

| Veto holder | Respondents | Share |
|---|---|---|
| Security, privacy, legal and compliance | 34 | 44% |
| Cross-functional business and technology governance | 33 | 43% |
| Finance, procurement, board and executive budget authority | 10 | 13% |

*Base: 77 respondents.*

### Which requirements have blocked adoption?

Of the 55 leaders who described a requirement that slowed or blocked adoption, 38% named security, retention, audit and compliance requirements; the rest split evenly between legacy-system integration and single sign-on (SSO), permissions and provisioning friction.

> "The SSO has been a huge problem for us. Because of how people were being added as users, there's a process where I have to go through and add some people in certain situations, and there's a time lag." — Department model selection, education

### When is switching providers justified?

When the gains are meaningful. Of the 49 leaders who weighed switching or admin overhead, 55% said meaningful performance, cost or security gains justify a switch, and a third said incumbent ecosystems, contracts and licensing keep them where they are.

### Does customer-facing use raise the bar?

Yes. Of 58 leaders, 45% said customer-facing and sensitive workflows raise the accuracy and safety bar. Of 81 leaders, 33% have built custom monitoring, review and workflow guardrails on top of vendor controls, and 42% rely on built-in access, privacy and retention controls.

**Strategic implication:** Sell to the veto. Publish access controls, retention guarantees, audit logs and SSO as a checklist mapped to the buyer's security review, and make provisioning painless.

**TL;DR:** Security, legal and governance teams decide LLM deployment, and the controls that matter are access, audit, retention and SSO.

---

## Are non-US LLMs such as DeepSeek and Qwen getting into the enterprise?

Rarely, and not because they failed a test. Of the 101 leaders asked about models from non-US labs such as DeepSeek, Kimi and Qwen, 75% said they are excluded or have never been tested in their organization. Only 27 of the 102 interviewed described running one on a real workload.

| Lived evaluation of non-US models | Respondents | Share |
|---|---|---|
| Described running a non-US model on a workload | 27 | 26% |
| Have not run one | 75 | 74% |

*Base: 102 interviews.*

### Why are non-US models excluded?

Data residency, IP and compliance, which 62% of 93 leaders called hard constraints; a quarter named geopolitical and procurement restrictions. A municipal finance and procurement director "ended up not being able to go with them" because state and federal rules bar contracting with technology companies originating in China. Of the 90 who have not run one, 60% cited security, privacy and geopolitical skepticism.

> "I've heard of these products. My initial instinct to answer your question is interesting, and I like to see what they have. But so far policy has been non US, no go." — IT, manufacturing

### What did buyers who tested non-US models find?

Among the 27 with lived evaluation, DeepSeek, Kimi and Qwen could match US models on simple, well-bounded tasks at lower cost. The finding is directional: across the full sample, only 5% concluded the cost case holds for simple workloads.

### Would a cheaper US efficiency tier change this?

Not on its own. Only 15% of 93 leaders expressed conditional interest in a US frontier lab launching an efficiency tier priced close to non-US models.

### Do regulated and technology buyers differ?

Yes, sharply. Regulated and industrial buyers exclude non-US models at 80%; technology and services buyers at 61%.

| Criterion | Regulated and industrial | Technology and services | Financial services (directional) |
|---|---|---|---|
| Exclude or never tested non-US models | 80% (36 of 45) | 61% (19 of 31) | 81% (13 of 16) |
| Security or data exposure was the dealbreaker | 62% (20 of 32) | 52% (13 of 25) | 55% (6 of 11) |
| Security, legal or compliance holds the veto | 42% (15 of 36) | 48% (12 of 25) | 30% (3 of 10) |
| Have a task-level cost-per-outcome unit | 24% (10 of 42) | 50% (15 of 30) | 6% (1 of 16) |
| Run a multi-provider environment | 50% (23 of 46) | 90% (28 of 31) | 53% (8 of 15) |
| Expect fragmentation to win | 61% (22 of 36) | 42% (10 of 24) | 75% (9 of 12) |

*Regulated and industrial: healthcare, education, public sector, industrial. Technology and services: software, technology, professional services. Segments under 30 are directional.*

**Strategic implication:** Answer the residency question before the price question. US labs should publish residency, retention and hosting guarantees first, and treat regulated and technology buyers as two different markets.

**TL;DR:** Three in four enterprises keep non-US models out on residency and compliance grounds that a lower price does not address.

---

## How do enterprises route work across LLM providers, and what triggers a switch?

Manually, by task. Of 101 leaders, 63% run multi-provider estates spanning several frontier labs, a fifth are evaluating a second provider alongside an incumbent, and a sixth run OpenAI or Microsoft Copilot alone. Of the 84 who described their routing logic, 80% assign work to models by hand.

| Routing logic | Respondents | Share |
|---|---|---|
| Manual task- and capability-based routing | 67 | 80% |
| Limited formal routers, emerging in-house gateways and platform routing | 10 | 12% |
| Cost, quality, latency and data-sensitivity routing | 7 | 8% |

*Base: 84 respondents.*

### What triggers an LLM provider switch?

Performance. Of the 84 leaders who described their last model change, 57% said it followed performance or capability, a quarter followed an IT, engineering or executive decision, and 15% followed cost. Portability is a live strategy for nearly half of the 72 who discussed it.

| Switch (from → to) | Trigger | Respondent |
|---|---|---|
| Cursor → Claude | Paywall added mid-contract | Program Management, technology |
| Cursor → Claude Code | Judged the better dev platform | CIO and CTO, healthcare technology |
| Copilot → Claude | Generated better, cheaper per result | PMO Director, technology consulting |
| ChatGPT → Claude | Better coding quality | AI Governance, pharmaceutical research |
| ChatGPT → Claude | Did the job; CFO cut a redundant tool | SVP of Marketing, software |
| ChatGPT → Gemini (testing) | Better answers on agent prompts | Lead Product Manager, retail e-commerce |
| ChatGPT → Copilot | Privacy, security and ERP fit | Logistics Director, supply chain |
| Anthropic → ChatGPT | New release handled context better | Senior Director of Strategy, technology and agency |
| AWS model → Salesforce model | Data already ran through Tableau | SVP for IT Strategy and Innovation, financial services |
| Gemini → other providers | "Not stable at the moment" | CEO, technology |
| Opus / Sonnet → cheaper tier (~$2) | Same value, far lower price | AI Productivity and AI DLC Initiative, enterprise SaaS |
| Lovable → in-house tool on Claude | Vendor data breach | Head of a division, big tech |

*Twelve individual examples selected from the 84 respondents who described their last model change; not a count. Nine were triggered by capability and fit, two by commercial terms, one by a security incident.*

### Do enterprises assign different jobs to different providers?

Yes: 90% of 79 leaders described provider specialization by task and workflow, for example one lab for coding and reasoning, another for general research. Of 88 who matched a model to a use case, 75% rely on use-case-specific capability and workflow fit.

> "Different AI models has different strength. We use Perplexity for article research and complex portfolio modeling and simulations. And ChatGPT for more general research or if you just want to get a quick answer or doing comparison chart or run certain Excel functions." — Technology Committee, investment and portfolio management

### Which providers do buyers trust?

Of 94 leaders, 49% trust established US and incumbent ecosystem providers most; most of the rest say security, data controls and regulatory alignment shape trust. Among the 29 who revised an opinion of a lab in the last six months, 18 had become more favorable toward Anthropic's Claude (small and directional).

### Will enterprises use more LLM providers or fewer?

More. 51% of 88 expect broader multi-provider and specialized model use over the next twelve months, and 55% of 80 chose fragmentation over consolidation.

| Expected direction (next 12 months) | Respondents | Share |
|---|---|---|
| Fragmentation through specialized multi-provider portfolios | 44 | 55% |
| Consolidation around durable incumbent providers | 34 | 42% |
| Named a market signal instead of a direction | 2 | 2% |

*Base: 80 respondents.*

**G2 perspective:** "63% of enterprises already run models from several labs, yet 80% of those who described their routing still decide by hand which model gets which job. That's a product opportunity hiding in plain sight. Buyers don't need another model. They need a simpler way to put the right one on the right task." — Praveen Maloo, Senior Director, AI Product Management, G2

**Strategic implication:** Build for the portfolio buyers already run. Publish which jobs a model is best at in the buyer's terms, make it easy to add alongside an incumbent, and make a switch cheap to test. For routing and gateway vendors, the market is the 80% who still route by hand.

**TL;DR:** Multi-model estates are the norm, routing is manual, and switches follow performance gaps, not leaderboards.

---

## What do buyers say in their own words?

> "We actually used public benchmarks to select to do an initial selection of which models we were going to use. And, yes, Microsoft's Copilot did mislead us, and the results we got when we applied real workloads was very different from what we expected." — Head of Information Security and AI Acceleration, footwear and apparel

> "We are moving from the culture of use AI for everything to use AI where it makes sense because of the cost of the company using AI solutions. The bill has gone very high very quickly as we push that out." — Senior Director of Managed Cloud Delivery, IT managed services

> "So the concern is IP for sure. Most important, data residency is the second one. So these two are absolutely nonnegotiable." — Content Director, pharmaceutical

> "DeepSeek was tested on customer support chat logs, and it seriously demonstrated a massive lower cost and fast response times comparable to the top US models." — AI Specialist, enterprise software

*Quotes are verbatim with light transcript cleanup. Participant organizations are anonymized.*

---

## What should LLM providers and vendors do?

1. **Publish workload evidence, not leaderboard positions (critical).** 85% of buyers treat benchmarks as background or let their own tests decide. Ship reference evaluations for the workloads buyers name, with cost and failure modes attached, and offer the pilot before the pitch.
2. **Price and report by outcome, and hide the token (critical).** 53% of buyers have no cost-per-outcome metric. A published cost per resolved ticket, document or case, with a forecastable ceiling, gives finance the number it needs to approve scale.
3. **Make every model change visible and reversible (high).** 40% of buyers have been surprised by a regression or silent model change. Version pinning, dated changelogs and regression suites on the buyer's own prompts make predictability visible.
4. **Lead with the security review (high).** Security or data exposure was the near-dealbreaker for 59% of buyers who named one. Map controls to the buyer's checklist and ship audit, retention and SSO on day one.
5. **Answer residency before price on the non-US question (high).** 75% of buyers exclude non-US models on residency, IP and compliance grounds. Publish hosting, retention and residency guarantees first.

**TL;DR:** Test it, price it by outcome and keep it predictable.

---

## FAQ

### Do enterprises actually use LLM benchmarks to choose a model?
Rarely as the deciding factor. Among the 93 leaders who described where benchmarks sat in their last model decision, 44% called them background context or ignored them and 41% said real-workload testing was the decisive evidence. Asked how much benchmark discourse helps, 55% of 92 said it has limited or background value and 42% prefer pilots and direct testing.

### How do enterprises evaluate LLMs before buying?
By testing them on their own workloads. 61% of 100 leaders are involved in hands-on model evaluation and selection. The criteria weighed most were task capability, output quality and accuracy (42% of 96), then security, privacy and enterprise integration (29%) and cost, speed and hands-on workflow testing (29%).

### Is cost per token the right way to measure LLM spend?
Buyers say no, but few have replaced it. 60% of 98 leaders compare models on downstream ROI and labor savings rather than price per million tokens, and only 23% track aggregate spend and token efficiency. Yet 53% of 96 have no formal cost-per-outcome metric; 29% measure per case, ticket, document or task and 18% count time saved.

### What blocks enterprises from adopting a new LLM?
Security and data exposure. 59% of the 75 leaders who named a near-dealbreaker cited security, privacy or data-exposure risk, ahead of failure on real-world task quality (28%) and unfavorable pricing (13%). Security, privacy, legal and compliance teams hold the veto for 44% of 77.

### Who approves LLM purchases in large companies?
Usually a shared group, with security or governance teams holding the veto. Decision rights are shared for 45% of 100 leaders, and the veto sits with security, privacy, legal or compliance for 44% of 77 and with a cross-functional governance group for 43%.

### Do LLM updates break production systems?
Sometimes, and quietly. 40% of 97 leaders said their last bad surprise in production was an output regression, hallucination or silent model change. Asked directly about provider updates, 21% of 78 said silent updates or version changes caused regressions, and 8% have testing, monitoring, rollback or fallback protections in place.

### Are DeepSeek and other non-US LLMs replacing US models in enterprises?
No, not in this sample. 75% of 101 leaders said non-US models such as DeepSeek, Kimi and Qwen are excluded or have never been tested, and 93% have limited or no comparative production evidence. Only 27 of the 102 interviewed have run one on a real workload.

### What percentage of companies use more than one LLM provider?
63% of 101 enterprises interviewed run multi-provider environments spanning several frontier labs. A further 20% are evaluating a second provider alongside their current one, and 17% run OpenAI or Microsoft Copilot alone. Routing is manual for 80% of the 84 who described it.

### Why do companies switch LLM providers?
Performance first, cost second. Among the 84 leaders who described their last model change, 57% said it was driven by performance or capability, 27% by an IT, engineering, business or executive decision, and 15% by cost, pricing or redundant spend.

---

## Methodology

This research draws on 102 in-depth interviews with enterprise leaders across the United States involved in choosing LLMs or the budget for model spend, conducted in August and September 2026 as conversational, AI-led, open-ended interviews. Interviews ran up to 34 minutes, with an average of 23 minutes.

**Who was interviewed.** Senior technology, finance, security and procurement leaders, including CEOs, CIOs, CTOs, CFOs, finance directors, heads of AI and information security, and directors of IT and engineering. 61% are hands-on in model evaluation and selection, 21% sit in a cross-functional governance or recommendation role and 18% own the budget or approve spend. Of the 93 who named an industry, 49% work in healthcare, education, public sector and industrial organizations, 33% in technology, software and professional services and 17% in financial services, insurance and banking. Company names are withheld.

**How it was analyzed.** Each interview response was coded to one of three or four labels per question by AI semantic analysis of the transcript, followed by multi-iteration validation and cross-verification, and every transcript was independently reviewed by G2's AI Custom Research team. Percentages are calculated on the respondents who addressed each question, and the base is stated with every figure. Quotations are verbatim, with light cleanup of transcription stutters and misheard product names only.

**Limits.** Bases range from 28 to 101. Figures resting on fewer than 30 people, such as the 27 who have run a non-US model, are directional. All respondents are in the United States and most are hands-on evaluators, so the findings describe these buyers rather than the whole market.

---

*Source: G2 Research — "What CFOs, CIOs and CTOs Weigh When Choosing an LLM" (G2 AI Custom Research, September 29, 2026; updated October 5, 2026). Based on G2 first-party research: 102 in-depth interviews with US enterprise LLM decision-makers. Full report: [What CFOs, CIOs and CTOs Weigh When Choosing an LLM](https://research-hub.g2.com/enterprise-llm-buying). Related: [AI Pricing: Proof Before Premium](https://research-hub.g2.com/ai-pricing-proof-before-premium) and the [G2 Large Language Models category](https://www.g2.com/categories/large-language-models-llms). Markdown edition structured for answer-engine optimization (AEO).*
