AI Insights

The top 10 LLMs in production

Who makes them, what they cost, and what the ten most-used LLM families actually run in production today.

  • ChatGPT and Gemini hold about 84% of consumer chat traffic, while Anthropic leads enterprise API spend
  • Coding is the largest paid workload, with Claude holding an estimated 54% of enterprise AI coding
  • Chinese open-weight models carry a majority of tokens routed through OpenRouter
  • Six models sit within three points on the public intelligence index
  • Production teams now pair closed frontier models with open or budget models for high-volume work
By Pradeep Chandran10 min read
Abstract intelligence themed top ten ranking with luminous model nodes arranged around a central neural list on a navy background.

Research report | 30 August 2026

Who makes them, what they cost, and what they actually run today

Ranked by real usage, consumer scale, enterprise API spend and routed token share, not by benchmark placement. Every figure is dated and sourced, current to 27 August 2026.

1B+ChatGPT weekly active users, Jul 2026
1BGemini app monthly users, Aug 2026
40%Enterprise API spend on Claude, Dec 2025
13%Enterprise workloads on open weights

1. What the numbers say

This report ranks the ten most-used large language model families as of August 2026 and documents what each one runs in production. Ordering follows usage rather than benchmark placement, blending consumer reach, enterprise spend and developer token volume. Where the underlying sources disagree, section 5 reports the disagreement.

The headline of 2026 is that there is no single market anymore. Consumer assistants, enterprise APIs and developer-built agents have different leaders, different economics and different winners, and a ranking that ignores the split misleads on all three.

01

The market split in two. ChatGPT and Gemini hold about 84% of consumer chat traffic, while Anthropic leads enterprise API spend at 40% (Menlo Ventures, Dec 2025).

02

Coding is the largest paid workload. Claude holds an estimated 54% of enterprise AI coding, and Claude Code alone passed 10% of public GitHub commits.

03

Chinese open-weight models (DeepSeek, Qwen, Kimi, GLM) carry a majority of tokens routed through OpenRouter, but only 13% of enterprise workloads.

04

The frontier is crowded: six models sit within three points on the public intelligence index, so price, speed and deployment terms decide most picks.

05

Meta reversed course. Llama is in maintenance mode and the closed Muse family now powers Meta AI across its apps.

  • Consumer reach: weekly and monthly active users of the assistant each family powers (OpenAI, Google, Similarweb).
  • Enterprise spend: share of LLM API spend and of coding workloads (Menlo Ventures, Dec 2025).
  • Developer volume: share of tokens routed through OpenRouter as of mid-2026.

2. The top 10 at a glance

Ordered by production footprint, following the three signals above.

Rank Model Maker Weights Used for today
1GPT-5.6OpenAIClosedConsumer assistance at 1B-user scale
2Claude 5AnthropicClosedEnterprise coding agents, legal, finance
3Gemini 3.7GoogleClosedMultimodal work, Workspace, volume
4Muse SparkMetaClosedMeta AI across Meta's consumer apps
5DeepSeek V4DeepSeekOpen, MITLow-cost reasoning, self-hosted volume
6Qwen 3.8AlibabaSplitMultilingual work, 90,000+ enterprises
7Grok 4.6xAIClosedLive X and web data, news research
8GLM-5.3ZhipuAPI todayBudget coding, code security review
9Kimi K3Moonshot AIOpen, customSelf-hosted agentic coding runs
10MistralMistral AISplitEU data residency, small models

Weights column: whether the model can be downloaded and self-hosted, or is rented through an API.

3. The models in production

Prices are API rates per million tokens, input and output, re-checked against Artificial Analysis on 27 August 2026. The intelligence index is that firm's 0 to 100 composite.

3.1

GPT-5.6 family | OpenAI

ChatGPT passed 1 billion weekly active users on 31 July 2026, and OpenAI's annualized revenue reached $40 billion in August, with enterprise revenue overtaking consumer for the first time in July. The current generation ships in three tiers: Sol at the frontier, Terra as the everyday default, and Luna for high volume at $0.05 per measured task. In production this family is the general-purpose layer: consumer assistance at population scale, the largest third-party app ecosystem, and coding through Codex. It holds 27% of enterprise LLM API spend, second to Anthropic.

Closed weights | Sol $4 / $20 per 1M tokens | intelligence index 61

3.2

Claude 5 family | Anthropic

Claude holds 40% of enterprise LLM API spend against OpenAI's 27%, and an estimated 54% of the enterprise AI coding market as of mid-2026. Opus 5 has led the public intelligence rankings since late July. Production use concentrates where output carries risk: coding agents, with Claude Code alone passing 10% of public GitHub commits, plus legal and financial analysis and long-document work. The consumer footprint is small, roughly 2% of chatbot web traffic, which is why the enterprise numbers surprise people.

Closed weights | Opus 5 $5 / $25 per 1M tokens | intelligence index 63

3.3

Gemini 3 family | Google

The Gemini app crossed 1 billion monthly active users on 11 August 2026, and Gemini-powered AI Overviews reach about 2.5 billion people a month. Gemini 3.7 Flash, released 13 August, is the value pick for multimodal work: a million-token context and text, image, audio and video input at a promotional rate through December. Google holds 21% of enterprise API spend, and a common production pattern tiers Flash-Lite for routing, Flash for main calls and Pro for escalation.

Closed weights | 3.7 Flash $0.75 / $3.75 per 1M tokens, promotional | index 56

3.4

Muse Spark | Meta

Muse Spark 1.2 powers the free Meta AI across WhatsApp, Instagram and Facebook, which arguably makes it the most widely deployed model on this list. It also marks Meta's reversal on open weights: the Muse family from Meta Superintelligence Labs is closed, and Llama is in maintenance mode. The model reads a million tokens and accepts every input type including audio. One caution for confidential work: the discounted contributor tier licenses your prompts for training. Meta has kept a small open line alongside it, led by the 30-billion-parameter Muse Glimmer under Apache 2.0.

Closed weights | Spark 1.2 $1.25 / $4.25 per 1M tokens | index 57

3.5

DeepSeek V4 | DeepSeek

DeepSeek made near-frontier quality cheap and downloadable, and it started the shift that now defines the open ecosystem: Chinese open-weight models went from negligible to a majority of tokens routed through OpenRouter between late 2024 and mid-2026. V4 ships under MIT in two flavors, with Flash measured at $0.11 per finished task. Production use is high-volume routine work, self-hosted deployments and cost-sensitive coding, most often as the budget tier under a closed frontier model.

Open weights, MIT | Pro $1.32 / $3.96, Flash $0.44 / $1.32 | index 52 to 53

3.6

Qwen 3.8 | Alibaba

Qwen is the most-adopted open base, with reported uptake by more than 90,000 enterprises and the broadest family of deployable sizes, from on-device models to the flagship class. In 2026 the flagship split in two: the proprietary multimodal Qwen3.8-Max API and a downloadable 2.4-trillion-parameter text-only checkpoint under a commercially restricted license. Production strengths are multilingual work, Alibaba Cloud deployment and fine-tuning as a base for derivatives. Specialized variants cover coding, vision and audio, which is why so many downstream models start from a Qwen checkpoint.

Split: closed Max API, restricted open checkpoint | $2 / $6 per 1M tokens | index 58

3.7

Grok 4.6 | xAI

Grok is wired into X and the live web, and that remains the reason to pick it: research and monitoring that need this morning's information in the same call as the reasoning. The 12 August release scores level with the frontier chasing pack, but it is cheap per token and dear per task: $2 / $6 on the rate card against a measured $1.23 per finished task, more than GPT-5.6 Sol. Its context stops at 500,000 tokens, half of what most of this list reads, and analysts still question its enterprise readiness beyond that live-data niche.

Closed weights | $2 / $6 per 1M tokens | index 60 | 500K context

3.8

GLM-5.3 | Zhipu

GLM is the cost story of 2026. The 18 August release matches Kimi K3's score at $1.40 / $4.40 and finishes tasks cheaper than Grok 4.6 or GPT-5.6 Sol; Zhipu also reports it narrowly ahead of Anthropic's and OpenAI's frontier models at spotting real vulnerabilities in source code. The 5.3 weights are not released, so today it is an API. The older GLM-5.2 is open, and GLM-5.3-Flash ships MIT weights at $0.09 per measured task. Production use is high-volume coding and agent steps where the budget, not the ceiling, is the constraint.

API only today; 5.2 and 5.3-Flash are open | $1.40 / $4.40 per 1M tokens | index 60

3.9

Kimi K3 | Moonshot AI

Kimi K3 is the strongest model you can download today: weights out since late July under Moonshot's own license, a million-token context, text and image input, and some of the best published terminal and coding-agent results. It is built for agentic software engineering, where the model plans, writes, tests and fixes across many steps. Self-hosting is a real commitment, since 2.8 trillion total parameters need roughly 1.6 TB of weights, and the license is Moonshot's own rather than MIT, so reselling access needs a read.

Open weights, Moonshot license | $3 / $15 per 1M tokens | index 60

3.10

Mistral | Mistral AI

Mistral is Europe's flagship lab and the default answer to data residency: EU-hosted inference, GDPR-native handling and Apache 2.0 licensing across much of the range, which keeps legal review short. The 2026 lineup centers on Large 3 and Medium 3.5 plus small, efficient models that are cheap to run. Production use concentrates in EU enterprises, regulated industries and on-device work rather than at the benchmark frontier.

Open and API mix, Apache 2.0 on much of the range | EU-hosted inference

4. Open weights vs closed

Closed frontier

Closed means prompts travel to the provider. That can rule it out for some regulated workloads, but it also gives teams the highest ceiling for the hard tenth of the work.

Open or budget volume

Open models are dramatically cheaper for routine work, and free beyond hardware when self-hosted. They also let teams change hosts when price, terms, or latency shift.

The camps flipped in 2026. Meta, the lab that made open weights mainstream, went closed with Muse, while the strong downloadable models now come from Moonshot, DeepSeek, Zhipu and Alibaba's separate open checkpoint. The two markets diverged with them: open-weight models carry 13% of enterprise workloads and falling, yet a majority of tokens routed through OpenRouter, where developer-built agents and pipelines run, now lands on Chinese open weights.

  • Where does the data go? Closed means prompts travel to the provider, which rules it out for some regulated workloads.
  • What does volume cost? Open models are dramatically cheaper for routine work, and free beyond hardware when self-hosted.
  • How locked in are you? One provider means inheriting its price changes and deprecations; open weights allow a change of host.

The pattern most production teams have settled on uses both: a closed frontier model for the hard tenth of the work and open or budget models for the high-volume rest. Agents make the split matter more, because they burn far more tokens than chat does, and measured cost per task rather than the per-token rate card is the number that predicts the invoice.

5. Sources and caveats

The figures in this report come from sources with different methods and reporting dates. Four caveats apply to how they should be read.

  • Enterprise share figures come from Menlo Ventures' December 2025 survey. They are the most-cited estimate available, not a census.
  • Revenue figures are run-rate and unaudited. Reports put Anthropic at $74.1B and OpenAI at $41.3B annualized in July 2026, both second-hand.
  • User counts mix weekly and monthly actives. OpenAI reports weekly figures only; trackers disagree on when 1 billion was crossed, citing either the 31 July confirmation or February's official 900 million.
  • Intelligence index scores within three points are effectively ties. Six models sit inside three points at the top of the board.

Source list

  1. Artificial Analysis leaderboard, via MindsHub model comparison27 Aug 2026 | mindshub.ai
  2. Menlo Ventures, State of Generative AI in the EnterpriseDec 2025 | menlovc.com
  3. OpenAI weekly active user disclosures, compiledFeb-Jul 2026 | getairefs.com
  4. Reuters via Sensor Tower, ChatGPT app crosses 1B monthly usersJun 2026 | demandsage.com
  5. OpenRouter token routing analysis, China's open-weight shareJun 2026 | datagravity.dev
  6. Enterprise API and coding share analysis (Menlo, The Information)Jun 2026 | valueaddvc.com
  7. LLM Stats composite leaderboardAug 2026 | llm-stats.com
  8. Shakudo, top language models roundupAug 2026 | shakudo.io
  9. Wavect, open-weight LLM comparison and licenses7 Aug 2026 | wavect.io

Build grounded agents

Build agents that reason inside your business logic

See how lowtouch.ai turns enterprise rules, policies, and semantic context into governed agents running inside your appliance.

About the Author

Pradeep Chandran

Pradeep Chandran

Lead - Agentic AI & DevOps

Pradeep Chandran is a seasoned technology leader and a key contributor at lowtouch.ai, a platform dedicated to empowering enterprises with no-code AI solutions. With a strong background in software engineering, cloud architecture, and AI-driven automation, he is committed to helping businesses streamline operations and achieve scalability through innovative technology. At lowtouch.ai, Pradeep focuses on designing and implementing intelligent agents that automate workflows, enhance operational efficiency, and ensure data privacy. His expertise lies in bridging the gap between complex IT systems and user-friendly solutions, enabling organizations to adopt AI seamlessly. Passionate about driving digital transformation, Pradeep is dedicated to creating tools that are intuitive, secure, and tailored to meet the unique needs of enterprises.

LinkedIn →