Who makes them, what they cost, and what the ten most-used LLM families actually run in production today.

Research report | 30 August 2026
Who makes them, what they cost, and what they actually run today
Ranked by real usage, consumer scale, enterprise API spend and routed token share, not by benchmark placement. Every figure is dated and sourced, current to 27 August 2026.
This report ranks the ten most-used large language model families as of August 2026 and documents what each one runs in production. Ordering follows usage rather than benchmark placement, blending consumer reach, enterprise spend and developer token volume. Where the underlying sources disagree, section 5 reports the disagreement.
The headline of 2026 is that there is no single market anymore. Consumer assistants, enterprise APIs and developer-built agents have different leaders, different economics and different winners, and a ranking that ignores the split misleads on all three.
The market split in two. ChatGPT and Gemini hold about 84% of consumer chat traffic, while Anthropic leads enterprise API spend at 40% (Menlo Ventures, Dec 2025).
Coding is the largest paid workload. Claude holds an estimated 54% of enterprise AI coding, and Claude Code alone passed 10% of public GitHub commits.
Chinese open-weight models (DeepSeek, Qwen, Kimi, GLM) carry a majority of tokens routed through OpenRouter, but only 13% of enterprise workloads.
The frontier is crowded: six models sit within three points on the public intelligence index, so price, speed and deployment terms decide most picks.
Meta reversed course. Llama is in maintenance mode and the closed Muse family now powers Meta AI across its apps.
Ordered by production footprint, following the three signals above.
| Rank | Model | Maker | Weights | Used for today |
|---|---|---|---|---|
| 1 | GPT-5.6 | OpenAI | Closed | Consumer assistance at 1B-user scale |
| 2 | Claude 5 | Anthropic | Closed | Enterprise coding agents, legal, finance |
| 3 | Gemini 3.7 | Closed | Multimodal work, Workspace, volume | |
| 4 | Muse Spark | Meta | Closed | Meta AI across Meta's consumer apps |
| 5 | DeepSeek V4 | DeepSeek | Open, MIT | Low-cost reasoning, self-hosted volume |
| 6 | Qwen 3.8 | Alibaba | Split | Multilingual work, 90,000+ enterprises |
| 7 | Grok 4.6 | xAI | Closed | Live X and web data, news research |
| 8 | GLM-5.3 | Zhipu | API today | Budget coding, code security review |
| 9 | Kimi K3 | Moonshot AI | Open, custom | Self-hosted agentic coding runs |
| 10 | Mistral | Mistral AI | Split | EU data residency, small models |
Weights column: whether the model can be downloaded and self-hosted, or is rented through an API.
Prices are API rates per million tokens, input and output, re-checked against Artificial Analysis on 27 August 2026. The intelligence index is that firm's 0 to 100 composite.
ChatGPT passed 1 billion weekly active users on 31 July 2026, and OpenAI's annualized revenue reached $40 billion in August, with enterprise revenue overtaking consumer for the first time in July. The current generation ships in three tiers: Sol at the frontier, Terra as the everyday default, and Luna for high volume at $0.05 per measured task. In production this family is the general-purpose layer: consumer assistance at population scale, the largest third-party app ecosystem, and coding through Codex. It holds 27% of enterprise LLM API spend, second to Anthropic.
Claude holds 40% of enterprise LLM API spend against OpenAI's 27%, and an estimated 54% of the enterprise AI coding market as of mid-2026. Opus 5 has led the public intelligence rankings since late July. Production use concentrates where output carries risk: coding agents, with Claude Code alone passing 10% of public GitHub commits, plus legal and financial analysis and long-document work. The consumer footprint is small, roughly 2% of chatbot web traffic, which is why the enterprise numbers surprise people.
The Gemini app crossed 1 billion monthly active users on 11 August 2026, and Gemini-powered AI Overviews reach about 2.5 billion people a month. Gemini 3.7 Flash, released 13 August, is the value pick for multimodal work: a million-token context and text, image, audio and video input at a promotional rate through December. Google holds 21% of enterprise API spend, and a common production pattern tiers Flash-Lite for routing, Flash for main calls and Pro for escalation.
Muse Spark 1.2 powers the free Meta AI across WhatsApp, Instagram and Facebook, which arguably makes it the most widely deployed model on this list. It also marks Meta's reversal on open weights: the Muse family from Meta Superintelligence Labs is closed, and Llama is in maintenance mode. The model reads a million tokens and accepts every input type including audio. One caution for confidential work: the discounted contributor tier licenses your prompts for training. Meta has kept a small open line alongside it, led by the 30-billion-parameter Muse Glimmer under Apache 2.0.
DeepSeek made near-frontier quality cheap and downloadable, and it started the shift that now defines the open ecosystem: Chinese open-weight models went from negligible to a majority of tokens routed through OpenRouter between late 2024 and mid-2026. V4 ships under MIT in two flavors, with Flash measured at $0.11 per finished task. Production use is high-volume routine work, self-hosted deployments and cost-sensitive coding, most often as the budget tier under a closed frontier model.
Qwen is the most-adopted open base, with reported uptake by more than 90,000 enterprises and the broadest family of deployable sizes, from on-device models to the flagship class. In 2026 the flagship split in two: the proprietary multimodal Qwen3.8-Max API and a downloadable 2.4-trillion-parameter text-only checkpoint under a commercially restricted license. Production strengths are multilingual work, Alibaba Cloud deployment and fine-tuning as a base for derivatives. Specialized variants cover coding, vision and audio, which is why so many downstream models start from a Qwen checkpoint.
Grok is wired into X and the live web, and that remains the reason to pick it: research and monitoring that need this morning's information in the same call as the reasoning. The 12 August release scores level with the frontier chasing pack, but it is cheap per token and dear per task: $2 / $6 on the rate card against a measured $1.23 per finished task, more than GPT-5.6 Sol. Its context stops at 500,000 tokens, half of what most of this list reads, and analysts still question its enterprise readiness beyond that live-data niche.
GLM is the cost story of 2026. The 18 August release matches Kimi K3's score at $1.40 / $4.40 and finishes tasks cheaper than Grok 4.6 or GPT-5.6 Sol; Zhipu also reports it narrowly ahead of Anthropic's and OpenAI's frontier models at spotting real vulnerabilities in source code. The 5.3 weights are not released, so today it is an API. The older GLM-5.2 is open, and GLM-5.3-Flash ships MIT weights at $0.09 per measured task. Production use is high-volume coding and agent steps where the budget, not the ceiling, is the constraint.
Kimi K3 is the strongest model you can download today: weights out since late July under Moonshot's own license, a million-token context, text and image input, and some of the best published terminal and coding-agent results. It is built for agentic software engineering, where the model plans, writes, tests and fixes across many steps. Self-hosting is a real commitment, since 2.8 trillion total parameters need roughly 1.6 TB of weights, and the license is Moonshot's own rather than MIT, so reselling access needs a read.
Mistral is Europe's flagship lab and the default answer to data residency: EU-hosted inference, GDPR-native handling and Apache 2.0 licensing across much of the range, which keeps legal review short. The 2026 lineup centers on Large 3 and Medium 3.5 plus small, efficient models that are cheap to run. Production use concentrates in EU enterprises, regulated industries and on-device work rather than at the benchmark frontier.
Closed means prompts travel to the provider. That can rule it out for some regulated workloads, but it also gives teams the highest ceiling for the hard tenth of the work.
Open models are dramatically cheaper for routine work, and free beyond hardware when self-hosted. They also let teams change hosts when price, terms, or latency shift.
The camps flipped in 2026. Meta, the lab that made open weights mainstream, went closed with Muse, while the strong downloadable models now come from Moonshot, DeepSeek, Zhipu and Alibaba's separate open checkpoint. The two markets diverged with them: open-weight models carry 13% of enterprise workloads and falling, yet a majority of tokens routed through OpenRouter, where developer-built agents and pipelines run, now lands on Chinese open weights.
The pattern most production teams have settled on uses both: a closed frontier model for the hard tenth of the work and open or budget models for the high-volume rest. Agents make the split matter more, because they burn far more tokens than chat does, and measured cost per task rather than the per-token rate card is the number that predicts the invoice.
The figures in this report come from sources with different methods and reporting dates. Four caveats apply to how they should be read.
Build grounded agents
See how lowtouch.ai turns enterprise rules, policies, and semantic context into governed agents running inside your appliance.
About the Author

Pradeep Chandran
Lead - Agentic AI & DevOps
Pradeep Chandran is a seasoned technology leader and a key contributor at lowtouch.ai, a platform dedicated to empowering enterprises with no-code AI solutions. With a strong background in software engineering, cloud architecture, and AI-driven automation, he is committed to helping businesses streamline operations and achieve scalability through innovative technology. At lowtouch.ai, Pradeep focuses on designing and implementing intelligent agents that automate workflows, enhance operational efficiency, and ensure data privacy. His expertise lies in bridging the gap between complex IT systems and user-friendly solutions, enabling organizations to adopt AI seamlessly. Passionate about driving digital transformation, Pradeep is dedicated to creating tools that are intuitive, secure, and tailored to meet the unique needs of enterprises.