| A | B | C | D | E | F | G | H | I | J | K | L | M | N | O | P | Q | R | S | T | U | V | W | X | Y | Z | AA | AB | AC | AD | AE | AF | AG | AH | AI | AJ | AK | AL | AM | AN | AO | AP | AQ | AR | AS | AT | AU | ||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1 | (910 models) | Permalink: | https://lifearchitect.ai/models-table/ | Upgrade to Models Table Pro to unlock all columns: | https://lifearchitect.ai/models-table-pro/ | The Memo: | https://lifearchitect.ai/memo | Filter using column arrows, or use Menu > Data > Create filter view. | Timeline view: | https://lifearchitect.ai/timeline | Compute | Hopper | Blackwell | Rubin | More... | |||||||||||||||||||||||||||||||||
2 | Model | Lab | Playground | Params (total, B) | Params (active, B) | Arch | Tokens trained (B) | Estimated total params? | Data ratio (total) | 🔒Training cost ($) | ALScore | MMLU | MMLU -Pro | GPQA | HLE | Training dataset | Announced ▼ | Public? | 🔒License | 🔒Context window | Disclosure score | Paper / Repo | Tags | Notes | Count (rough) | 🔒Params total confidence | 🔒Params active confidence | 🔒Tokens confidence | 🔒Country | 🔒Training hardware | 🔒Compute (FLOPs) | 🔒Compute (Log FLOPs) | 🔒Compute (ZettaFLOPs) | 🔒Frontier compute % | 🔒Compute percentile | 🔒H100 hours | 🔒H100 energy (MWh) | 🔒H100 CO2 Emissions (tonnes) | 🔒B200 hours | 🔒B200 cost to train | 🔒B200 energy (MWh) | 🔒B200 CO2 Emissions (tonnes) | 🔒Rubin hours (hold) | 🔒Rubin cost to train (hold) | 🔒Rubin energy (MWh) | 🔒Rubin CO2 Emissions (tonnes) | 🔒More private columns | |
3 | Muse Spark 1.1 | Meta AI | https://meta.ai/ | 500 | 50 | MoE | 120,000 | * | 240:1 | 25.8 | 62.1 | synthetic, web-scale | Jul/2026 | 🟢 | D | https://ai.meta.com/static-resource/muse-spark-1-1-evaluation-report | Reasoning, SOTA | Multimodal reasoning model from Meta Superintelligence Labs. First model on Meta Model API ($1.25/$4.25 per 1M tokens). Trained for agentic tasks with active 1M-token context management; multi-agent orchestration and thought compression. MSL Research: "Muse Spark 1.1 is a significant upgrade of 1.0, especially on general STEM reasoning, coding, personal and professional agentic tasks. We incorporated more and higher quality data, spent significantly more human research compute and GPU compute with a more stable async RL stack." https://x.com/shuchaobi/status/2075235778561798462 | 910 | |||||||||||||||||||||||||||||
4 | Grok 4.5 | xAI | https://grok.com/ | 1500 | 75 | MoE | 120,000 | 80:1 | 44.7 | synthetic, web-scale | Jul/2026 | 🟢 | C | https://x.ai/news/grok-4-5 | Reasoning | First model built jointly with Cursor. V9 foundation architecture; 1.5T params; MoE. Trained on tens of thousands of NVIDIA GB300 GPUs. 80 TPS; $2/$6 per M tokens. Announce: https://cursor.com/blog/grok-4-5 | 909 | |||||||||||||||||||||||||||||||
5 | Horus-Hiero-9B | TokenAI | https://huggingface.co/tokenaii/Horus-Hiero-9B | 9 | Dense | 40,000 | 4,445:1 | 2.0 | 79.3 | 78.1 | synthetic, web-scale | Jul/2026 | 🟢 | C | https://huggingface.co/tokenaii/Horus-Hiero-9B | Reasoning | Fine-tune of Qwen3.5-9B for hieroglyph translation and 150+ languages. First hieroglyphic-capable LLM from Egyptian AI startup. Multimodal (text/image/video). | 908 | ||||||||||||||||||||||||||||||
6 | Hy3 | Tencent | https://huggingface.co/tencent/Hy3 | 295 | 21 | MoE | 40,000 | 136:1 | 11.5 | 90.4 | 53.2 | synthetic, web-scale | Jul/2026 | 🟢 | C | https://github.com/Tencent-Hunyuan/Hy3 | Reasoning | Post-training upgrade of Hy3 Preview after feedback from 50+ product teams. 192 experts (top-8 activated) with 3.8B MTP layer. Hallucination rate reduced from 12.5% to 5.4%. Scored 2.67/4 vs GLM-5.1 at 2.51/4 in 312 blind expert comparisons. SWE-bench Verified accuracy variance within 4% across scaffoldings (CodeBuddy/Cline/KiloCode). License changed from Community to Apache 2.0. | 907 | |||||||||||||||||||||||||||||
7 | Leanstral 1.5 | Mistral | https://huggingface.co/mistralai/Leanstral-1.5-119B-A6B | 119 | 6.5 | MoE | 15,000 | 127:1 | 4.5 | synthetic, web-scale | Jul/2026 | 🟢 | C | https://mistral.ai/news/leanstral-1-5/ | Reasoning | Open-source code agent model for Lean 4 theorem proving only (text-only), part of the Mistral Small 4 family (128 experts, 4 active). Saturates miniF2F (100%), solves 587/672 PutnamBench, SOTA on FATE-H (87%) and FATE-X (34%); found 5 previously unknown bugs across 57 real repositories. | 906 | |||||||||||||||||||||||||||||||
8 | Hierarchos 232M | Independent | https://github.com/necat101/Hierarchos | 0.232 | Dense | 0.13 | 1:1 | 0.001 | web-scale | Jul/2026 | 🟢 | C | https://github.com/necat101/Hierarchos/blob/main/HIERARCHOS_FINDINGS_PAPER.md | Experimental hybrid non-Transformer (RWKV v8 + Titans memory + HRM). 13 epochs on Alpaca SFT on RTX 6000 Blackwell. 'Dataset: netcat420/Experiment_0.1 (Alpaca format)' + 'Training: 13 epochs'. Standard Alpaca ~52K samples × ~200 tokens/sample × 13 epochs ≈ 135M. Smoke-test (n=100): ARC Easy 0.36, HellaSwag 0.37, TruthfulQA 0.22. Announce: https://www.reddit.com/r/MachineLearning/comments/1um123n/hierarchos_preliminary_findings_from_a_232m/ | 905 | |||||||||||||||||||||||||||||||||
9 | Laguna XS 2.1 | Poolside | https://huggingface.co/poolside/Laguna-XS-2.1 | 33.4 | 3 | MoE | 30,000 | 899:1 | 3.3 | 53 | synthetic, web-scale | Jul/2026 | 🟢 | A | https://poolside.ai/assets/laguna/laguna-m1-xs2-technical-report.pdf | Reasoning | 33B total / 3B active MoE for agentic coding on a local machine; upgrade of XS.2 with +5.4pt SWE-bench Multilingual. SWE-bench Verified 70.9, SWE-bench Multilingual 63.1, SWE-Bench Pro 47.6, Terminal-Bench 2.0 37.5. MMLU-Pro from XS.2 base model report. Announce: https://poolside.ai/blog/introducing-laguna-xs-2-1 | 904 | ||||||||||||||||||||||||||||||
10 | Nemotron-Labs-TwoTower-30B-A3B | NVIDIA | https://huggingface.co/nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16 | 60 | 6 | MoE | 27,100 | 452:1 | 4.3 | 78.24 | 60.93 | synthetic, web-scale | Jul/2026 | 🟢 | A | https://arxiv.org/abs/2606.26493 | Diffusion | Two-tower block-wise autoregressive diffusion LLM built on Nemotron-3-Nano-30B-A3B. Frozen AR context tower + trainable diffusion denoiser tower. Retains 98.7% of AR baseline quality at 2.42x generation throughput. 128K context. NVIDIA Nemotron Open Model License. | 903 | |||||||||||||||||||||||||||||
11 | TabFM 1.0 | https://huggingface.co/google/tabfm-1.0.0-pytorch | 1.6 | Dense | 600 | * | 375:1 | 0.1 | synthetic | Jun/2026 | 🟢 | D | https://github.com/google-research/tabfm | Zero-shot tabular foundation model for classification and regression via in-context learning. Trained entirely on hundreds of millions of synthetic datasets. Blog: "trained entirely on hundreds of millions of synthetic datasets." Comparable models: TabPFN v2 trained on 130M datasets; TabICLv2 on ~33M+ datasets across 550K steps. "Hundreds of millions" → ~300M datasets. TabICLv2's curriculum uses datasets from 1K to 60K rows; average ~2K. So ~300M × 2K rows = ~600B tabular data points. This parallels TimesFM's "100B time-points" convention for non-text foundation models. #1 on TabArena. Being integrated into Google BigQuery. | 902 | |||||||||||||||||||||||||||||||||
12 | Claude Sonnet 5 | Anthropic | https://claude.ai/ | 1000 | 20 | MoE | 80,000 | * | 80:1 | 29.8 | 89 | 57.4 | synthetic, web-scale | Jun/2026 | 🟢 | D | https://www-cdn.anthropic.com/d9bb04416ffe1352af84721476c1fa9994c07fde/Claude%20Sonnet%205%20System%20Card.pdf | Reasoning | 1M context. Announce: https://www.anthropic.com/news/claude-sonnet-5. Showing GMMLU (Global MMLU by Cohere). | 901 | ||||||||||||||||||||||||||||
13 | openPangu-2.0-Flash | Huawei | https://huggingface.co/openpangu/openPangu-2.0-Flash | 92 | 6 | MoE | 34,000 | 370:1 | 5.9 | 83.7 | synthetic, web-scale | Jun/2026 | 🟢 | A | https://ai.gitcode.com/ascend-tribe/openPangu-2.0-Flash | Reasoning | First frontier-scale MoE model trained entirely on non-NVIDIA hardware (Ascend 910B NPUs). 512K context. MLA + DSA/SWA hybrid attention (1:2 ratio). 3-head MTP (Multi-Token Prediction). Muon optimizer. Post-training: SFT + multi-objective RL + online policy distillation (OPD). GPQA-Diamond=83.7 (Thinking; Avg@4). AIME 2026=93.3 (Avg@16). SWE-bench Verified=63.1 (Avg@3). License: openPangu Model License Agreement v2.0. | 900 | ||||||||||||||||||||||||||||||
14 | LongCat-2.0 | Meituan | https://longcat.ai | 1600 | 48 | MoE | 35,000 | 22:1 | 24.9 | 88.9 | synthetic, web-scale | Jun/2026 | 🟢 | A | https://longcat.chat/blog/longcat-2.0/ | Reasoning | Trained entirely on domestic AI ASIC superpods with no NVIDIA GPUs. LongCat Sparse Attention for 1M context. MOPD post-training from agent, reasoning, and interaction expert groups. | 899 | ||||||||||||||||||||||||||||||
15 | Agents-A1 | Shanghai AI Lab | https://huggingface.co/InternScience/Agents-A1 | 35 | 3 | MoE | 40,000 | 1,143:1 | 3.9 | 47.6 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://arxiv.org/abs/2606.30616 | Reasoning | 35B MoE agentic model fine-tuned from Qwen3.5-35B-A3B via three-stage recipe: full-domain SFT, domain-level teacher training, and multi-teacher on-policy distillation. Matches or outperforms 1T-parameter models on long-horizon agent benchmarks. | 898 | ||||||||||||||||||||||||||||||
16 | GPT-5.6 Sol | OpenAI | https://chatgpt.com/ | 3000 | 150 | MoE | 200,000 | * | 67:1 | 81.6 | synthetic, web-scale | Jun/2026 | 🟢 | D | https://deploymentsafety.openai.com/gpt-5-6-preview | Reasoning, SOTA | Flagship of GPT-5.6 family (Sol/Terra/Luna). SOTA Terminal-Bench 2.1. CTF saturated at 96.7%. $5/$30 per 1M tokens. Limited preview to trusted partners; broad release planned. Developer log analysis and pre-release reports suggest a 1.5M token context window, not yet officially confirmed by OpenAI in the blog or system card. Max and Ultra reasoning modes. 'We will share an expanded suite of evaluation results when we make the model broadly available.' | 897 | ||||||||||||||||||||||||||||||
17 | Ornith-1.0-397B | DeepReinforce | https://huggingface.co/collections/deepreinforce-ai/ornith-10 | 397 | 17 | MoE | 36,000 | 91:1 | 12.6 | synthetic, web-scale | Jun/2026 | 🟢 | A | https://deep-reinforce.com/ornith_1_0.html | Reasoning | Self-improving RL framework for agentic coding. Post-trained on Qwen 3.5-397B-A17B. 82.4 SWE-Bench Verified, 77.5 Terminal-Bench 2.1. | 896 | |||||||||||||||||||||||||||||||
18 | Unlimited-OCR | Baidu | https://huggingface.co/baidu/Unlimited-OCR | 3 | 0.5 | MoE | 6,000 | 2,000:1 | 0.4 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://arxiv.org/abs/2606.23050 | End-to-end OCR model building on DeepSeek-OCR; replaces all decoder attention with Reference Sliding Window Attention (R-SWA) for constant KV cache; 3B total / 500M active MoE; scores 93.23 on OmniDocBench v1.5 (+6% over DeepSeek-OCR); parses dozens of pages in a single forward pass at 32K max length. | 895 | ||||||||||||||||||||||||||||||||
19 | Seed2.1 Pro | ByteDance | https://exp.volcengine.com/ark?csid=excs-202602141507-%5BJk-cTyC-D-fHC56fiUG_K%5D&mode=chat&modelId=doubao-seed-2-1-pro-260628 | 500 | 50 | MoE | 30,000 | * | 60:1 | 12.9 | 55.7 | synthetic, web-scale | Jun/2026 | 🟢 | D | https://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/seed2.1/Seed2_1_Model_Card.pdf | Reasoning | Agentic productivity model. Ranked 8th on Code Arena: Frontend (1539). Compared against Claude Opus 4.7 and GPT-5.5. | 894 | |||||||||||||||||||||||||||||
20 | QUEST-35B-RL | OSU NLP | https://huggingface.co/osunlp/QUEST-35B-RL | 35 | 3 | MoE | 40,000 | 1,143:1 | 3.9 | 37.9 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://arxiv.org/abs/2605.24218 | Qwen3.5-35B-A3B base, Open deep research agent family (2B–35B) trained with fully synthetic rubric-tree tasks on 32 H100s. Approaches or surpasses frontier closed-source agents across eight deep research benchmarks. Announce: https://x.com/ysu_nlp/status/2067380438134624742 | 893 | |||||||||||||||||||||||||||||||
21 | Glimmer-1-Base | Glint Research | https://huggingface.co/Glint-Research/Glimmer-1-Base | 0.0000119 | Dense | 0.0005 | 43:1 | 0.000 | web-scale | Jun/2026 | 🟢 | A | https://huggingface.co/Glint-Research/Glimmer-1-Base | 11.9K-parameter (0.0000119B) experimental micro-model trained on 500K tokens of FineWeb-Edu. Llama-style transformer exploring the lower bound of useful language model scale. Base only, no SFT. Trained on FineWeb-Edu on a single RTX 4070 SUPER. | 892 | |||||||||||||||||||||||||||||||||
22 | GLM-5.2 | Z.AI | https://huggingface.co/zai-org/GLM-5.2 | 744 | 40 | MoE | 28,500 | 39:1 | 15.3 | 91.2 | 54.7 | synthetic, web-scale | Jun/2026 | 🟢 | A | https://arxiv.org/abs/2602.15763 | Reasoning | 1M-token context (up from 200K in GLM-5.1); 131K max output. Trained entirely on Huawei Ascend 910B; no NVIDIA hardware. Two thinking modes (High; Max). | 891 | |||||||||||||||||||||||||||||
23 | VibeThinker-3B | WeiboAI (Sina Weibo) | https://huggingface.co/WeiboAI/VibeThinker-3B | 3 | Dense | 5,500 | 1,834:1 | 0.4 | 70.2 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://arxiv.org/abs/2606.16140 | Reasoning | 3B dense reasoning model scoring 94.3 on AIME26 and 80.2 on LiveCodeBench v6; matches 100x+ larger models on verifiable reasoning tasks. Built on Qwen2.5-Coder-3B via curriculum SFT + multi-domain RL + offline self-distillation. | 890 | |||||||||||||||||||||||||||||||
24 | SubQ 1.1 Small | Subquadratic | https://subq.ai/request-early-access | 70 | Dense | 13,000 | * | 186:1 | 3.2 | 85.4 | synthetic, web-scale | Jun/2026 | 🟢 | D | https://subq.ai/docs/subq-1-1-small-model-card.pdf | Reasoning | SSA (Subquadratic Sparse Attention). Near-perfect NIAH retrieval to 12M tokens. GPQA Diamond 85.4%. LiveCodeBench v6 pass@4=89.7%. RULER 128K=99.12%. AutomationBench Finance=13%. Third-party verified by Appen. 64.5x less compute than dense attention at 1M tokens. | 889 | ||||||||||||||||||||||||||||||
25 | Rio-3.5-Open-397B | IplanRIO | https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B | 397 | 17 | MoE | 36,000 | 91:1 | 12.6 | 88 | 90.9 | 36.5 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B | Reasoning | Merge of Qwen 3.5 397B + Next-N2-Pro. IplanRIO is Rio de Janeiro municipal IT company. Features SwiReasoning: dynamic latent/explicit reasoning via entropy-based confidence signals. SWE-Bench Verified=80.2. https://github.com/nex-agi/Nex-N2/issues/4 | 888 | ||||||||||||||||||||||||||||
26 | openPangu-2.0-Pro | Huawei | pending 30/jun | 505 | 18 | MoE | 19,000 | 38:1 | 10.3 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://gitcode.com/ascend-tribe | Open-source MoE with record 28:1 sparsity ratio. DSA+SWA hybrid attention architecture. Optimized for Ascend NPU; 2x single-card throughput vs mainstream open-source models. Open-sourcing 7 components from 30/Jun/2026. Dataset: Predecessor openPangu-Ultra-MoE-718B (718B/39B active) trained on ~19T tokens. | 887 | ||||||||||||||||||||||||||||||||
27 | Kimi-K2.7-Code | Moonshot AI | https://huggingface.co/moonshotai/Kimi-K2.7-Code | 1000 | 32 | MoE | 30,500 | 31:1 | 18.4 | synthetic, web-scale | Jun/2026 | 🟢 | A | https://huggingface.co/moonshotai/Kimi-K2.7-Code | Reasoning | Coding-focused agentic model built upon Kimi K2.6. Reduces thinking-token usage by ~30% compared to K2.6. 15.5T is verified base pretraining only; K2.5→K2.6→K2.7 continued training adds undisclosed tokens | 886 | |||||||||||||||||||||||||||||||
28 | Nex-N2-Pro | Nex AGI | https://huggingface.co/nex-agi/Nex-N2-Pro | 397 | 17 | MoE | 36,000 | 91:1 | 12.6 | 90.7 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://github.com/nex-agi/Nex-N2 | Reasoning | Post-trained on Qwen3.5-397B-A17B. "An agentic model with Agentic Thinking." GPQA Diamond up from 88.4 (base Qwen3.5) to 90.7 (+2.3 from post-training). SWE-Bench Verified=80.8, Terminal-Bench 2.1=75.3, SWE-Bench Pro=58.8. Competitive with GPT-5.5 and Opus 4.7 on coding and agentic benchmarks. | 885 | ||||||||||||||||||||||||||||||
29 | DiffusionGemma 26B A4B IT | Google DeepMind | https://huggingface.co/google/diffusiongemma-26B-A4B-it | 25.2 | 3.8 | MoE | 14,000 | 556:1 | 2.0 | 77.6 | 73.2 | 11.9 | web-scale | Jun/2026 | 🟢 | C | https://huggingface.co/google/diffusiongemma-26B-A4B-it | Reasoning, Diffusion | Experimental discrete diffusion model on Gemma 4 26B A4B MoE backbone. Generates 256-token blocks in parallel via iterative denoising (1000+ tok/s on H100, 700+ on RTX 5090). Bidirectional attention enables self-correction. Quality lower than standard Gemma 4; designed for speed-critical local workflows. | 884 | ||||||||||||||||||||||||||||
30 | Apodex-1.0-H | Apodex AI | https://apodex.ai | 397 | 17 | MoE | 36,000 | * | 91:1 | 12.6 | 60.8 | synthetic, web-scale | Jun/2026 | 🟢 | D | https://www.apodex.com/blog/apodex-1.0 | Reasoning | Verification-centric deep-research agent team on Qwen3.5 base. Heavy-duty mode coordinates up to 150 sub-agents over 15,000 steps. SOTA on BrowseComp (90.3), DeepSearchQA (94.4), FrontierScience-Research (46.7). | 883 | |||||||||||||||||||||||||||||
31 | Claude Fable 5 | Anthropic | https://claude.ai/ | 10000 | 150 | MoE | 250,000 | * | 25:1 | 166.7 | 94.1 | 64.5 | synthetic, web-scale | Jun/2026 | 🟢 | D | https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf | Reasoning, SOTA | Mythos-class model made safe for general use. Same underlying model as Claude Mythos 5 with safety classifiers (fallback to Opus 4.8 in <5% of sessions for cyber, bio/chem, distillation). Pricing $10/$50 per Mtok. API: claude-fable-5. | 882 | ||||||||||||||||||||||||||||
32 | North-Mini-Code-1.0 | Cohere | https://huggingface.co/CohereLabs/North-Mini-Code-1.0 | 30 | 3 | MoE | 12,000 | 400:1 | 2.0 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://huggingface.co/blog/CohereLabs/introducing-north-mini-code | 30B-A3B MoE (128 experts, 8 active per token) optimized for agentic software engineering. First model in Cohere's North family. Artificial Analysis Coding Index: 33.4. 256K context, 64K output. | 881 | ||||||||||||||||||||||||||||||||
33 | AFM 3 Core Advanced | Apple | https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models | 20 | 4 | MoE | 25,000 | 1,250:1 | 2.4 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models | Most powerful Apple on-device model. 20B params stored in flash (NAND); 1–4B activated per prompt via Instruction-Following Pruning (IFP). Natively multimodal (text, image, audio). Built with Google on cloud TPUs. 2026 blog states 'we significantly scaled pre-training on the latest generation of cloud TPU accelerators' and 'all models shared a common initial foundation.' Conservative 15T estimate accounts for scaling over 14T+ base + multimodal tokens. | 880 | ||||||||||||||||||||||||||||||||
34 | AFM 3 Cloud Pro | Apple | https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models | 1200 | 60 | MoE | 60,000 | * | 50:1 | 28.3 | synthetic, web-scale | Jun/2026 | 🟢 | D | https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models | Apple–Google–NVIDIA collaboration. Based on custom 1.2T-parameter Gemini model with Apple's own pre-training and post-training. Runs on NVIDIA GPUs in Google Cloud via extended Private Cloud Compute. 'Our most capable server-based model, which powers our most demanding use cases, like agentic tool use and complex reasoning.' Tech report planned for summer 2026. Bloomberg, Mark Gurman, Nov 5 2025: "Apple Inc. is planning to pay about $1 billion a year for an ultrapowerful 1.2 trillion parameter artificial intelligence model developed by Alphabet Inc.'s Google" URI: https://www.bloomberg.com/news/articles/2025-11-05/apple-plans-to-use-1-2-trillion-parameter-google-gemini-model-to-power-new-siri "Based on Gemini foundation… Apple did their own pre-training, post-training" — Max Weinbach tweet, Jun 8 2026, quoted in wccftech: "Apple just clarified AFM Cloud is Apple's own model, trained with Gemini outputs / AFM local models are entirely Apple models / AFM Cloud Pro seems to be based on Gemini foundation and data, but Apple did their own pre-training, post-training, RL, etc" URI: https://wccftech.com/apple-removes-the-fog-around-its-new-cloud-based-and-20-billion-parameter-on-device-ai-models-brushes-aside-googles-contributions-while-hyping-nvidias/. | 879 | |||||||||||||||||||||||||||||||
35 | Macaron-V1-Preview-749B | Mind Lab | https://macaron-model-previews.macaron.im/ | 749 | 41 | MoE | 28,500 | 39:1 | 15.4 | synthetic, web-scale | Jun/2026 | 🟢 | A | https://macaron.im/mindlab/research/macaron-v1-preview | 749B Mixture-of-LoRA agent model post-trained from GLM-5.1 (744B frozen base + 5 × 1B specialist LoRAs for chat, personal-life, coding, Generative UI, and OpenClaw tasks). Router Tool routes between adapters. SWE-bench Verified=78.1. | 878 | ||||||||||||||||||||||||||||||||
36 | Gemma 4 12B | Google DeepMind | https://huggingface.co/google/gemma-4-12B-it | 12 | Dense | 14,000 | 1,167:1 | 1.4 | 77.2 | 78.8 | 5.2 | web-scale | Jun/2026 | 🟢 | C | https://arxiv.org/abs/2607.02770 | Reasoning | Encoder-free multimodal (text, image, audio) dense model with configurable thinking mode and 256K context. | 877 | |||||||||||||||||||||||||||||
37 | Aion-1.0-Plan | Microsoft | 14 | Dense | 14,000 | 1,000:1 | 1.5 | synthetic, web-scale | Jun/2026 | 🟢 | D | https://blogs.windows.com/windowsdeveloper/2026/06/02/build-2026-furthering-windows-as-the-trusted-platform-for-development/ | Reasoning | On-device reasoning and tool-calling SLM that ships in-box as part of Windows on capable devices, enabling fully local agentic workflows. "Enables applications to reason over user intent, invoke tools, manage files and orchestrate sub-agents, bringing fully agentic workflows onto the device." Announced at Build 2026; available in the coming months. | 876 | |||||||||||||||||||||||||||||||||
38 | Aion-1.0-Instruct | Microsoft | https://microsoftedge.github.io/Demos/built-in-ai/playgrounds/prompt-api/ | 2 | Dense | 8,000 | * | 4,000:1 | 0.4 | synthetic, web-scale | Jun/2026 | 🟢 | D | https://blogs.windows.com/msedgedev/2026/06/02/expanding-on-device-ai-in-microsoft-edge-new-models-and-apis-for-the-web/ | Pre-release small language model for on-device AI in Microsoft Edge (Canary/Dev), powering the Prompt and Writing Assistance APIs. Successor to Phi-4-mini (4B); "smaller, faster, and more efficient," supports CPU inference for devices without a GPU. Planned open-source release on Hugging Face in July 2026. | 875 | ||||||||||||||||||||||||||||||||
39 | MAI-Code-1-Flash | Microsoft | https://github.blog/changelog/2026-06-02-mai-code-1-flash-is-now-available-for-github-copilot/ | 30 | Dense | 15,000 | * | 500:1 | 2.2 | web-scale | Jun/2026 | 🟢 | D | https://microsoft.ai/news/introducingmai-code-1-flash/ | Reasoning | Lightweight agentic coding model from Microsoft AI, built end-to-end on clean and appropriately licensed data, trained directly with GitHub Copilot harnesses. Adaptive solution-length control: solves harder problems with up to 60% fewer tokens. Outperforms Claude Haiku 4.5 across SWE-Bench Verified, SWE-Bench Pro (51.2% vs 35.2%), SWE-Bench Multilingual, and Terminal Bench 2. Available in VS Code GitHub Copilot. | 874 | |||||||||||||||||||||||||||||||
40 | MAI-Thinking-1 | Microsoft | https://microsoft.ai/news/introducing-mai-thinking-1/ | 1000 | 35 | MoE | 33,500 | 34:1 | 19.3 | 85 | 84.2 | web-scale | Jun/2026 | 🟢 | A | https://microsoft.ai/wp-content/uploads/2026/06/main_20260602_2.pdf | Reasoning | Microsoft AI's reasoning model. 35B-active, ~1T-total parameters sparse MoE. Trained from the ground up without distillation from third-party models, on clean and commercially licensed data. Matches Claude Opus 4.6 on SWE-Bench Pro and preferred over Claude Sonnet 4.6 in blind human side-by-side evaluations. AIME 2025=97.0, AIME 2026=94.5. | 873 | |||||||||||||||||||||||||||||
41 | KeyLM-75M-Instruct | Independent | https://huggingface.co/Eclipse-Senpai/KeyLM-75M-Instruct | 0.0753 | Dense | 18 | 240:1 | 0.004 | 24 | web-scale | Jun/2026 | 🟢 | A | https://huggingface.co/Eclipse-Senpai/KeyLM-75M-Instruct | 75M-param from-scratch small LM; competitive on IFEval vs SmolLM-135M-Instruct at half the size. "trained completely on kaggle (tpu v5e-8)" Announce: https://www.reddit.com/r/LocalLLaMA/comments/1tuyb8s/i_trained_a_75m_parameter_llm_from_scratch_on_18b/ | 872 | ||||||||||||||||||||||||||||||||
42 | Cosmos 3 Super | NVIDIA | https://huggingface.co/nvidia/Cosmos3-Super | 64 | 32 | MoE | 200 | 4:1 | 0.4 | special | Jun/2026 | 🟢 | C | https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdf | SOTA | Omnimodal world model for Physical AI; dual-tower mixture-of-transformers (reasoner + generator) initialized from Qwen3-VL-32B. Dataset: ‘two epochs over the full pre-training mixture’ with sequences ‘at most 16k tokens.’ Conservative avg ~4K tokens/sample × 22M × 2 epochs ≈ 176B pretrain + ~9B SFT ≈ ~185B; rounded to 200B. Excludes generator-pathway vision/audio/action tokens (hundreds of millions of images and videos, not directly token-comparable to LLM corpora) and the prior Qwen3-VL-32B initialization tokens (not disclosed). | 871 | |||||||||||||||||||||||||||||||
43 | Mellum2-12B-A2.5B-Thinking | JetBrains | https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Thinking | 12 | 2.5 | MoE | 10,600 | 884:1 | 1.2 | 57.6 | synthetic, web-scale | Jun/2026 | 🟢 | A | https://arxiv.org/abs/2605.31268 | Reasoning | Open-weight 12B MoE (64 experts, 8 active) language model specialised in software engineering; successor to the 4B dense Mellum. | 870 | ||||||||||||||||||||||||||||||
44 | Qwen3.7-Plus | Alibaba | https://chat.qwen.ai/ | 480 | 35 | MoE | 40,000 | * | 84:1 | 14.6 | 88.5 | 90.3 | 34.7 | synthetic, web-scale | Jun/2026 | 🟢 | D | https://qwen.ai/blog?id=qwen3.7-plus | Reasoning | Multimodal agent model unifying vision and language; operates GUI and CLI within a single agent loop | 869 | |||||||||||||||||||||||||||
45 | Nemotron 3 Ultra | NVIDIA | https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16#nvidia-nemotron-3-ultra-550b-a55b-bf16 | 550 | 55 | MoE | 25,000 | 46:1 | 12.4 | 89.1 | 86.8 | 87 | 37.4 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://arxiv.org/abs/2512.20856 | NVIDIA’s largest open model: 550B total parameters with up to 55B active per token via a hybrid Mamba-Transformer MoE architecture. Most intelligent US open weights model per Artificial Analysis (Intelligence Index 48). | 868 | ||||||||||||||||||||||||||||
46 | MiniMax-M3 | MiniMax | https://huggingface.co/MiniMaxAI/MiniMax-M3 | 428 | 23 | MoE | 100,000 | 234:1 | 21.8 | synthetic, web-scale | Jun/2026 | 🟢 | C | https://www.minimax.io/blog/minimax-m3 | Reasoning, SOTA | "M3 is a model that has undergone mixed-modality training from Step 0... After rebuilding the entire data pipeline for this data, we are now able to scale the training data to the order of 100 trillion tokens." First open-weight model combining frontier coding, 1M-token MSA context, and native multimodality. SWE-Bench Pro: 59.0; BrowseComp: 83.5 (surpasses Opus 4.7 at 79.3). Tech report and weights to be released within 10 days of launch. | 867 | |||||||||||||||||||||||||||||||
47 | Step 3.7 Flash | StepFun | https://huggingface.co/stepfun-ai/Step-3.7-Flash | 198 | 11 | MoE | 24,000 | 122:1 | 7.3 | 49.7 | synthetic, web-scale | May/2026 | 🟢 | A | https://github.com/stepfun-ai/Step-3.7-Flash | Reasoning | A high-efficiency Flash model for real-world agents. | 866 | ||||||||||||||||||||||||||||||
48 | LFM2.5-8B-A1B | Liquid AI | https://huggingface.co/LiquidAI/LFM2.5-8B-A1B | 8.3 | 1.5 | MoE | 38,000 | 4,579:1 | 1.9 | synthetic, web-scale | May/2026 | 🟢 | A | https://www.liquid.ai/blog/lfm2-5-8b-a1b | Reasoning | Edge MoE for fast on-device tool calling; 128K context, reasoning-only model. Highlights: IFEval 91.84, MATH500 88.76, BFCLv3 64.79, Tau2-Telecom 88.07. | 865 | |||||||||||||||||||||||||||||||
49 | Claude Opus 4.8 | Anthropic | https://claude.ai/ | 5000 | 150 | MoE | 80,000 | * | 16:1 | 66.7 | 93.6 | 57.9 | synthetic, web-scale | May/2026 | 🟢 | D | https://www.anthropic.com/claude-opus-4-8-system-card | Reasoning, SOTA | Announce: https://www.anthropic.com/news/claude-opus-4-8 HLE=with tools (49.8 no tools). Same price as Opus 4.7 ($5/$25). Params/tokens carried from Opus-class estimate (Opus 4.7). | 864 | ||||||||||||||||||||||||||||
50 | ESMC 6B | Biohub | https://huggingface.co/biohub/ESMC-6B | 6 | Dense | 6,600 | 1,100:1 | 0.7 | special | May/2026 | 🟢 | A | https://biohub.ai/papers/esm_protein.pdf | “Language Modeling Materializes a World Model of Protein Biology”. Protein language model, 80 layers, 2.37e23 training FLOPs. Trained on ~2.8B protein sequences (UniRef + MGnify + JGI, clustered at 70% identity). Tokens back-calculated from disclosed compute: training FLOPs of 2.37e23 reported on the HF model card, divided by 6N (Kaplan/Chinchilla rule for transformer training), gives 2.37e23 ÷ (6 × 6e9) ≈ 6.6T tokens. Released alongside ESMFold2 and ESM Atlas (6.8B sequences / 1.1B predicted structures). MIT license. | 863 | |||||||||||||||||||||||||||||||||
51 | MiniCPM5-1B | OpenBMB | https://huggingface.co/spaces/openbmb/MiniCPM5-1B-Demo | 1.08 | Dense | 8,000 | 7,408:1 | 0.3 | 48.85 | 26.26 | synthetic, web-scale | May/2026 | 🟢 | A | https://huggingface.co/openbmb/MiniCPM5-1B | Reasoning | the first model in the MiniCPM5 series. It is a dense 1B Transformer built for on-device, local deployment, and resource-constrained scenarios, reaching 1B-class open-source SOTA. Hybrid reasoning with <think> template. 1,080,632,832 params, 24 layers, GQA 16Q/2KV, 131K context. Post-training: 200B deep-thinking SFT + 200B hybrid-thinking SFT + RL + OPD. Pretraining token count not disclosed; estimate aligned with MiniCPM4 paper's 8T-token series budget (UltraClean/Ultra-FineWeb data-efficient pipeline). MMLU-Redux=70.06, MATH-500=91.6, AIME-2025=40.42, BBH=71.89, IFEval=80.41 (Thinking mode). | 862 | ||||||||||||||||||||||||||||||
52 | Gated DeltaNet-2 | NVIDIA | https://github.com/NVlabs/GatedDeltaNet-2 | 1.3 | Dense | 100 | 77:1 | 0.04 | web-scale | May/2026 | 🟢 | A | https://github.com/NVlabs/GatedDeltaNet-2/blob/main/paper/GDN2_paper.pdf | Linear-attention architecture decoupling channel-wise erase and write gates; 1.3B params trained on 100B tokens of FineWeb-Edu; outperforms Mamba-2, Gated DeltaNet, KDA, and Mamba-3 on language modeling, RULER, and retrieval. | 861 | |||||||||||||||||||||||||||||||||
53 | Command A+ | Cohere | https://huggingface.co/CohereLabs/command-a-plus-05-2026-bf16 | 218 | 25 | MoE | 20,000 | 92:1 | 7.0 | synthetic, web-scale | May/2026 | 🟢 | C | https://huggingface.co/CohereLabs/command-a-plus-05-2026-bf16 | "open source model with 25 billion active parameters and 218B total parameters model optimized for agentic, multilingual, and reasoning-heavy tasks with a focus on enterprise performance, while also providing support for vision inputs". 128 experts, 8 active per token + 1 shared expert. 128K context. 48 languages. Announce: https://cohere.com/blog/command-a-plus | 860 | ||||||||||||||||||||||||||||||||
54 | Qwen3.7-Max | Alibaba | https://chat.qwen.ai/ | 2000 | 100 | MoE | 40,000 | * | 20:1 | 29.8 | 89.6 | 92.4 | 53.5 | synthetic, web-scale | May/2026 | 🟢 | D | https://qwen.ai/blog?id=qwen3.7 | Reasoning | "Qwen3.7-Max, our latest proprietary model designed for the agent era." 35-hour autonomous kernel optimization run with 1,000+ tool calls; 10.0x geomean speedup over Triton reference. Available soon via Alibaba Cloud Model Studio. | 859 | |||||||||||||||||||||||||||
55 | HRM-Text-1B | Sapient Intelligence | https://huggingface.co/sapientinc/HRM-Text-1B | 1 | Dense | 160 | 160:1 | 0.04 | 60.7 | synthetic, web-scale | May/2026 | 🟢 | A | https://github.com/sapientinc/HRM-Text | Reasoning | "1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning." Pretraining cost ~$1472 on 16 H100s. 40B unique tokens × 4 epochs = 160B total. | 858 | |||||||||||||||||||||||||||||||
56 | Nemotron-Labs-Diffusion-14B | NVIDIA | https://huggingface.co/nvidia/Nemotron-Labs-Diffusion-14B | 14 | Dense | 4,345 | * | 311:1 | 0.8 | 82.51 | 54.55 | synthetic, web-scale | May/2026 | 🟢 | D | https://d1qx31qr3h6wln.cloudfront.net/publications/Nemotron_Diffusion_Tech_Report_v1.pdf | Diffusion | Tri-mode LM unifying AR, diffusion, and self-speculation decoding within a single architecture. Explicit Nemotron training is disclosed at 1T (Stage 1 AR) + 300B (Stage 2 joint) + 45B (SFT) = 1,345B. Base init from Ministral3-14B (Liu et al., arXiv 2601.08584, 2026)≈3T (disclosed). | 857 | |||||||||||||||||||||||||||||
57 | Gemini 3.5 Flash | Google DeepMind | https://gemini.google.com/ | 500 | 25 | MoE | 100,000 | 200:1 | 23.6 | 40.2 | synthetic, web-scale | May/2026 | 🟢 | A | https://deepmind.google/models/model-cards/gemini-3-5-flash/ | “Gemini 3.5 Flash delivers intelligence that rivals large flagship models on multiple dimensions, at the speeds you have come to expect from the Flash series.” Beats Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), MCP Atlas (83.6%), CharXiv Reasoning (84.2%). Pro coming next month. Announce: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/ | 856 | |||||||||||||||||||||||||||||||
58 | ZAYA1-8B-Diffusion-Preview | Zyphra | https://huggingface.co/Zyphra/ZAYA1-8B | 8.4 | 0.76 | MoE | 15,100 | 1,798:1 | 1.2 | 94 | synthetic, web-scale | May/2026 | 🟢 | A | https://www.zyphra.com/post/zaya1-8b-diffusion-preview | Diffusion | First MoE diffusion model converted from an autoregressive LLM, and first diffusion-LM trained on AMD. Built from ZAYA1-8B base via TiDAR-style conversion: 600B tokens at 32k + 500B context extension to 128k + diffusion SFT. 4.6x decoding speedup (lossless sampler), 7.7x (mixed-logits sampler). Diffuses blocks of 16 tokens simultaneously. | 855 | ||||||||||||||||||||||||||||||
59 | Intern-S2-Preview | Shanghai AI Laboratory/SenseTime | https://huggingface.co/internlm/Intern-S2-Preview | 35 | 3 | MoE | 41,000 | 1,172:1 | 4.0 | 88 | 18.07 | synthetic, web-scale | May/2026 | 🟢 | A | https://huggingface.co/internlm/Intern-S2-Preview | Reasoning | "an efficient 35B scientific multimodal foundation model... continued pretrained from Qwen3.5." 35B-A3B MoE. Tokens estimate: Qwen3.5 base ~36T + ~5T scientific continued pretraining (matching Intern-S1-Pro pattern) ≈ 41T. Announce: https://github.com/InternLM/Intern-S1 | 854 | |||||||||||||||||||||||||||||
60 | Ring-2.6-1T | Inclusion AI | https://huggingface.co/inclusionAI/Ring-2.6-1T | 1000 | 50 | MoE | 20,750 | 21:1 | 15.2 | 88.27 | synthetic, web-scale | May/2026 | 🟢 | C | https://arxiv.org/abs/2606.15079 | Reasoning | Reasoning sibling of Ling-2.6-1T. Async RL training + IcePop algorithm. Two reasoning effort levels: high and xhigh. xhigh: AIME 26=95.83, ARC-AGI-V2=66.18, GPQA Diamond=88.27. high: PinchBench=87.60, ClawEval=63.82, Tau2-Bench Telecom=95.32. | 853 | ||||||||||||||||||||||||||||||
61 | TML-Interaction-Small | Thinking Machines Lab | 276 | 12 | MoE | 30,000 | 109:1 | 9.6 | web-scale, audio, video | May/2026 | 🟡 | D | https://thinkingmachines.ai/blog/interaction-models/ | Native multimodal (audio, video, text) interaction model with time-aligned 200ms micro-turns. “TML-Interaction-Small dominates interaction quality while being more intelligent than any non thinking model.” Research preview; larger models planned later in 2026. | 852 | |||||||||||||||||||||||||||||||||
62 | Needle | Cactus-Compute | https://huggingface.co/Cactus-Compute/needle | 0.026 | Dense | 202 | 7,770:1 | 0.008 | synthetic, web-scale | May/2026 | 🟢 | A | https://github.com/cactus-compute/needle | Distilled from Gemini 3.1 into a 26M-parameter "Simple Attention Network" (attention + gating, no MLPs/FFNs). 6000 tok/s prefill, 1200 tok/s decode on consumer devices. Pretrained on 16 TPU v6e for 200B tokens (27 hrs), post-trained on 2B tokens of single-shot function-calling data (45 min) synthesized via Gemini across 15 tool categories. MIT licensed. Weights: https://huggingface.co/Cactus-Compute/needle | 851 | |||||||||||||||||||||||||||||||||
63 | NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B | NVIDIA | https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B-BF16 | 30 | 3.6 | MoE | 25,160 | 839:1 | 2.9 | 78.63 | 72.1 | synthetic, web-scale | May/2026 | 🟢 | A | https://arxiv.org/abs/2511.16664 | "NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B-BF16 is a 3-in-1 elastic large language model (LLM) developed by NVIDIA. It contains three nested model variants (30B, 23B, and 12B parameters) within a single BF16 checkpoint, all sharing the same parameter space." | 850 | ||||||||||||||||||||||||||||||
64 | ZAYA1-74B-Preview | Zyphra | https://huggingface.co/Zyphra/ZAYA1-74B-preview | 74 | 4 | MoE | 18,000 | 244:1 | 3.8 | 84.4 | 83.8 | synthetic, web-scale | May/2026 | 🟢 | A | https://www.zyphra.com/zaya1-8b-technical-report | Reasoning | "Pre-RL reasoning base checkpoint (no instruction or RL post-training). Trained end-to-end on AMD MI300x hardware. Uses CCA attention with sliding window attention hybrid and 256k context. "ZAYA1-74B-Preview is a pre-RL reasoning-base checkpoint, released under an Apache 2.0 license." | 849 | |||||||||||||||||||||||||||||
65 | SubQ 1M-Preview | Subquadratic | https://subq.ai/request-early-access | 70 | Dense | 12,000 | * | 172:1 | 3.1 | synthetic, web-scale | May/2026 | 🟢 | D | https://subq.ai/how-ssa-makes-long-context-practical | First fully subquadratic LLM. SSA (Subquadratic Sparse Attention). 12M token research context, 1M production. SWE-Bench Verified=81.8. | 848 | ||||||||||||||||||||||||||||||||
66 | Llama 3.1 8B + CUA (quantum) | Multiverse Computing | 8 | Dense | 15,000 | 1,875:1 | 1.2 | synthetic, web-scale | May/2026 | 🟡 | B | https://arxiv.org/abs/2605.05914 | Quantum | First end-to-end quantum enhancement of a production-scale LLM on real superconducting hardware (156-qubit IBM Quantum System Two); Cayley unitary adapters add 6,000 params and reduce Llama 3.1 8B perplexity by 1.4%. | 847 | |||||||||||||||||||||||||||||||||
67 | ZAYA1-8B | Zyphra | https://huggingface.co/Zyphra/ZAYA1-8B | 8.4 | 0.76 | MoE | 14,000 | 1,667:1 | 1.1 | 74.2 | 71 | synthetic, web-scale | May/2026 | 🟢 | A | https://www.zyphra.com/zaya1-8b-technical-report | Reasoning | First MoE model pretrained, midtrained, and SFT’d entirely on AMD Instinct MI300; 760M active params; Apache-2.0. Announce: https://www.zyphra.com/post/zaya1-8b | 846 | |||||||||||||||||||||||||||||
68 | GENE-26.5 | Genesis AI | https://www.genesis.ai/blog/gene-26-5-advancing-robotic-manipulation-to-human-level | 30 | Dense | 10,000 | * | 334:1 | 1.8 | robotics | May/2026 | 🟡 | D | https://www.genesis.ai/blog/gene-26-5-advancing-robotic-manipulation-to-human-level | "our first robotic foundation model system and the initial public release in the GENE family… designed to push general-purpose robotic manipulation towards human-level capability." 200,000+ hours of multimodal data (vision, hand state, language, tactile, robot controls). 200,000 hours × ~2M tokens/hour ≈ 400B multimodal tokens, + internet language/video pretraining priors ≈ 10T total. | 845 | ||||||||||||||||||||||||||||||||
69 | GPT-5.5 Instant | OpenAI | https://chatgpt.com/ | 300 | 15 | MoE | 114,000 | * | 380:1 | 19.5 | 85.6 | synthetic, web-scale | May/2026 | 🟢 | D | https://openai.com/index/gpt-5-5-instant/ | Reasoning, SOTA | "smarter and more accurate, with clearer, more concise answers that feel better tailored to you. Because Instant is the daily driver for hundreds of millions of people..." | 844 | |||||||||||||||||||||||||||||
70 | MAMMAL | IBM | https://huggingface.co/ibm/biomed.omics.bl.sm.ma-ted-458m | 0.458 | Dense | 700 | 1,529:1 | 2.3 | special | May/2026 | 🟢 | A | https://www.nature.com/articles/s44386-026-00047-4 | "MAMMAL (Molecular Aligned Multi Modal Architecture and Language), a foundation model for cross-modal learning, designed to address the challenges associated with drug discovery tasks." 2 Billion total samples × ~350 average tokens/sample ≈ 700 Billion tokens. | 843 | |||||||||||||||||||||||||||||||||
71 | ERNIE-5.1-Preview | Baidu | https://ernie.baidu.com/ | 800 | 60 | MoE | 100,000 | 125:1 | 29.8 | 84.3 | 91 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://ernie.baidu.com/blog/posts/ernie-5.1-preview-0430-release-on-lmarena/ | Reasoning | "ERNIE-5.1-Preview builds on the pre-training foundation of ERNIE-5.0 while compressing total parameters to approximately 1/3 and active parameters to approximately 1/2, achieving leading performance at its model scale using only about 6% of the pre-training cost of comparable models." | 842 | |||||||||||||||||||||||||||||
72 | Granite-4.1-30B | IBM | https://huggingface.co/ibm-granite/granite-4.1-30b | 30 | Dense | 15,000 | 500:1 | 2.3 | 80.16 | 64.09 | 45.76 | synthetic, web-scale | Apr/2026 | 🟢 | A | https://huggingface.co/blog/ibm-granite/granite-4-1 | Reasoning | "Granite 4.1 is a family of dense, decoder‑only LLMs (3B, 8B, and 30B) trained on ~15T tokens using a multi‑stage pre‑training pipeline, including long‑context extension of up to 512K tokens." | 841 | |||||||||||||||||||||||||||||
73 | Mistral Medium 3.5 | Mistral | https://chat.mistral.ai/chat | 128 | Dense | 12,000 | 94:1 | 4.1 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5 | "a dense 128B model with a 256k context window, handling instruction-following, reasoning, and coding in a single set of weights" | 840 | |||||||||||||||||||||||||||||||||
74 | Nemotron 3 Nano Omni | NVIDIA | https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 | 30 | 3 | MoE | 25,700 | 857:1 | 2.9 | 77.3 | 72.2 | synthetic, web-scale | Apr/2026 | 🟢 | A | https://arxiv.org/abs/2604.24954 | Reasoning | Built on the highly efficient Nemotron 3 Nano 30B-A3B backbone, Nemotron 3 Nano Omni natively supports audio inputs alongside text, images, and video. ~717B tokens (multimodal post-training) on top of ~25T base LLM pretraining, ~25.7T total. | 839 | |||||||||||||||||||||||||||||
75 | Laguna XS.2 | Poolside | https://huggingface.co/poolside/Laguna-XS.2 | 33 | 3 | MoE | 10,000 | 304:1 | 1.9 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://poolside.ai/blog/introducing-laguna-xs2-m1 | "Laguna XS.2, as open-weights. 33B total parameters, 3B active. Apache 2.0." | 838 | ||||||||||||||||||||||||||||||||
76 | Laguna M.1 | Poolside | https://platform.poolside.ai/ | 225 | 23 | MoE | 10,000 | 45:1 | 5.0 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://poolside.ai/blog/introducing-laguna-xs2-m1 | "Laguna M.1 is a 225B total parameter model with 23B activated parameters, built for agentic coding and long-horizon work." | 837 | ||||||||||||||||||||||||||||||||
77 | DeepSeek-V4-Pro | DeepSeek-AI | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro | 1600 | 49 | MoE | 33,000 | 21:1 | 24.2 | 90.1 | 87.5 | 90.1 | 37.7 | synthetic, web-scale | Apr/2026 | 🟢 | A | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf | SOTA, Reasoning | MoE with 1.6T total / 49B active parameters, 1M-token context, trained on 32T+ tokens with FP4/FP8 mixed precision; introduces Compressed Sparse Attention (CSA) using only 27% single-token FLOPs vs. V3.2 and 10% KV cache; scores 90.1 on GPQA Diamond and 80.6% on SWE-bench Verified. Open-weight. | 836 | |||||||||||||||||||||||||||
78 | talkie-1930-13b | Independent | https://talkie-lm.com/chat | 13 | Dense | 260 | 20:1 | 0.2 | history only | Apr/2026 | 🟢 | A | https://talkie-lm.com/introducing-talkie | Alec Radford (GPT-1, GPT-2, GPT-3). talkie-1930-13b is a 13b language model trained on pre-1931 English-language text, instruction-tuned using a novel instruction-following dataset built from pre-1931 reference works including etiquette manuals, letter-writing manuals, encyclopedias, and poetry collections. It has also undergone reinforcement learning using online DPO to improve instruction-following capabilities. | 835 | |||||||||||||||||||||||||||||||||
79 | Hy3 preview | Tencent | https://huggingface.co/tencent/Hy3-preview | 295 | 21 | MoE | 40,000 | 136:1 | 11.5 | 87.42 | 65.76 | 30 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://github.com/Tencent-Hunyuan/Hy3-preview | Reasoning | "Hy3 preview is the first model trained on our rebuilt infrastructure, and the strongest we've shipped so far. It improves significantly on complex reasoning, instruction following, context learning, coding, and agent tasks." | 834 | ||||||||||||||||||||||||||||
80 | Ling-2.6-1T | Inclusion AI | https://huggingface.co/inclusionAI/Ling-2.6-1T | 1000 | 50 | MoE | 20,750 | 21:1 | 15.2 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://arxiv.org/abs/2606.15079 | SOTA | "excellence across reasoning, coding, and tool-calling, achieving open-source SOTA status on multiple execution-heavy benchmarks:" | 833 | |||||||||||||||||||||||||||||||
81 | GPT-5.5 | OpenAI | https://chatgpt.com/ | 3000 | 150 | MoE | 200,000 | * | 67:1 | 81.6 | 93.6 | 57.2 | synthetic, web-scale | Apr/2026 | 🟢 | D | https://deploymentsafety.openai.com/gpt-5-5/gpt-5-5.pdf | Reasoning, SOTA | Announce: https://openai.com/index/introducing-gpt-5-5/ HLE result is for GPT-5.5 Pro/x-high. | 832 | ||||||||||||||||||||||||||||
82 | Marul V7 | Independent | https://marulai.com.tr/ | 0.258 | Dense | 1,000 | 3,876:1 | 0.05 | web-scale | Apr/2026 | 🟢 | C | https://www.reddit.com/r/LocalLLaMA/comments/1sshwtu/s%C4%B1f%C4%B1rdan_e%C4%9Fitilmi%C5%9F_258m_parametre_t%C3%BCrk%C3%A7e_llm/ | Turkish. | 831 | |||||||||||||||||||||||||||||||||
83 | Qwen3.6-27B | Alibaba | https://huggingface.co/Qwen/Qwen3.6-27B | 27 | Dense | 40,000 | 1,482:1 | 3.5 | 86.1 | 87.8 | 24.3 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://huggingface.co/Qwen/Qwen3.6-27B | Reasoning | "the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience." | 830 | |||||||||||||||||||||||||||||
84 | MiMo-V2.5-Pro | Xiaomi | https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro | 1020 | 42 | MoE | 27,000 | 27:1 | 17.5 | 89.4 | 68.5 | 66.7 | 48 | synthetic, web-scale | Apr/2026 | 🟢 | A | https://mimo.xiaomi.com/mimo-v2-5-pro | Reasoning | Uses a 7:1 Hybrid Attention mechanism and supports a 1M-token context window. "significant improvements over its predecessor, MiMo-V2-Pro, in general agentic capabilities, complex software engineering, and long-horizon tasks." | 829 | |||||||||||||||||||||||||||
85 | Ling-2.6-Flash | Inclusion AI | https://huggingface.co/inclusionAI/Ling-2.6-flash | 104 | 7.4 | MoE | 20,750 | 200:1 | 4.9 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://arxiv.org/abs/2606.15079 | Reasoning | MoE with 104B total / 7.4B active params using a 1:7 MLA + Lightning Linear hybrid attention architecture for up to 4x throughput vs. comparable models; supports 262K context; achieves 61.2% on SWE-bench Verified and 73.85% on MathArena AIME 2026. Open-weight. | 828 | |||||||||||||||||||||||||||||||
86 | Granite-4.1-8B | IBM | https://huggingface.co/ibm-granite/granite-4.1-8b | 8 | Dense | 15,000 | 1,875:1 | 2.3 | 73.84 | 55.99 | 41.96 | synthetic, web-scale | Apr/2026 | 🟢 | A | https://huggingface.co/ibm-granite/granite-4.1-8b | Reasoning | "improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities." | 827 | |||||||||||||||||||||||||||||
87 | OpenMythos | Independent | https://github.com/kyegomez/OpenMythos | 0.77 | 0.04 | MoE | 30 | 39:1 | 18.4 | web-scale | Apr/2026 | 🟢 | A | https://github.com/kyegomez/OpenMythos | 770M trained, up to 1T available. "OpenMythos is an open-source, theoretical implementation of the Claude Mythos model. It implements a Recurrent-Depth Transformer (RDT) with three stages: Prelude (transformer blocks), a looped Recurrent Block (up to max_loop_iters), and a final Coda. Attention is switchable between MLA and GQA, and the feed-forward uses a sparse MoE with routed and shared experts ideal for exploring compute-adaptive, depth-variable reasoning." | 826 | ||||||||||||||||||||||||||||||||
88 | Kimi K2.6 | Moonshot AI | https://huggingface.co/moonshotai/Kimi-K2.6 | 1000 | 32 | MoE | 30,500 | 31:1 | 18.4 | 90.5 | 54 | synthetic, web-scale | Apr/2026 | 🟢 | A | https://www.kimi.com/blog/kimi-k2-6 | Reasoning, SOTA | "Kimi K2.6 is an open-source, native multimodal agentic model" | 825 | |||||||||||||||||||||||||||||
89 | Qwen3.6-Max-Preview | Alibaba | https://chat.qwen.ai/ | 2000 | 100 | MoE | 40,000 | * | 20:1 | 20.0 | synthetic, web-scale | Apr/2026 | 🟢 | D | https://qwen.ai/blog?id=qwen3.6-max-preview | Reasoning | "Qwen3.6-Max-Preview is an early preview of our next proprietary model, delivering meaningful improvements over Qwen3.6-Plus in agentic coding, world knowledge, and instruction following. It achieves the top score on six major coding benchmarks — SWE-bench Pro, Terminal-Bench 2.0, SkillsBench, QwenClawBench, QwenWebBench, and SciCode — with substantial gains over its predecessor. It also demonstrates stronger knowledge (SuperGPQA, QwenChineseBench) and better instruction following (ToolcallFormatIFBench)." | 824 | ||||||||||||||||||||||||||||||
90 | Qwen3.6-35B-A3B | Alibaba | https://huggingface.co/Qwen/Qwen3.6-35B-A3B | 35 | 3 | MoE | 40,000 | 1,143:1 | 3.9 | 85.2 | 86 | 21.4 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://qwen.ai/blog?id=qwen3.6-35b-a3b | Reasoning | "Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6." | 823 | ||||||||||||||||||||||||||||
91 | Grok 4.3 | xAI | https://grok.com/ | 500 | 25 | MoE | 80,000 | 160:1 | 21.1 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://grok.com/release-notes | Reasoning | "Grok 4.3 is a new pre-trained model matching the scale of Grok 4.20 with an improved architecture and a December 2025 knowledge cutoff." "0.5T total. Current Grok [4.2] is half the size of Sonnet and 1/10th the size of Opus." https://x.com/elonmusk/status/2042123561666855235 "The public facing v4.2 is based on foundation model v8, trained on Hoppers, with significant shortfalls in training data quality, comprehensiveness and proportionality. It is also only 0.5T in size." https://x.com/elonmusk/status/2055298325994164377 | 822 | |||||||||||||||||||||||||||||||
92 | Claude Opus 4.7 | Anthropic | https://claude.ai/ | 5000 | 150 | MoE | 80,000 | * | 16:1 | 66.7 | 94.2 | 54.7 | synthetic, web-scale | Apr/2026 | 🟢 | D | https://cdn.sanity.io/files/4zrzovbb/website/037f06850df7fbe871e206dad004c3db5fd50340.pdf | Reasoning, SOTA | Announce: https://www.anthropic.com/news/claude-opus-4-7 | 821 | ||||||||||||||||||||||||||||
93 | GPT-Rosalind | OpenAI | 3000 | 150 | MoE | 114,000 | * | 38:1 | 61.6 | synthetic, web-scale | Apr/2026 | 🔴 | F | https://openai.com/index/introducing-gpt-rosalind/ | Reasoning | "our frontier reasoning model built to support research across biology, drug discovery, and translational medicine. The life sciences model series is optimized for scientific workflows, combining improved tool use with deeper understanding across chemistry, protein engineering, and genomics." | 820 | |||||||||||||||||||||||||||||||
94 | GPT-5.4-Cyber | OpenAI | 3000 | 150 | MoE | 114,000 | * | 38:1 | 61.6 | synthetic, web-scale | Apr/2026 | 🔴 | F | https://openai.com/index/scaling-trusted-access-for-cyber-defense/ | Reasoning | "In preparation for increasingly more capable models from OpenAI over the next few months, we are fine-tuning our models specifically to enable defensive cybersecurity use cases, starting today with a variant of GPT‑5.4 trained to be cyber-permissive: GPT‑5.4‑Cyber. " | 819 | |||||||||||||||||||||||||||||||
95 | Marco-Mini | Alibaba | https://huggingface.co/AIDC-AI/Marco-Mini-Instruct | 17.3 | 0.86 | MoE | 36,000 | 2,081:1 | 2.6 | 83.4 | 70.7 | 50.3 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://huggingface.co/AIDC-AI/Marco-Mini-Instruct | Reasoning | "a highly sparse MoE multilingual model... 256 experts, 8 active per token. Drop-Upcycling from Qwen3-0.6B-Base. 2-stage post-training: SFT + Online Policy Distillation" Dataset: ""Web-mined NLLB translation pairs... Wikidata-sourced content with Gemini3-Flash text synthesis for cultural concepts." | 818 | ||||||||||||||||||||||||||||
96 | EXAONE 4.5 | LG | https://huggingface.co/LGAI-EXAONE/EXAONE-4.5-33B | 33 | Dense | 14,000 | 425:1 | 2.3 | 83.3 | 80.5 | 13.6 | web-scale | Apr/2026 | 🟢 | C | https://arxiv.org/abs/2604.08644 | Reasoning | “EXAONE”=“EXpert AI for EveryONE”. Training tokens/ratio: EXAONE-3 7.8B=8T tokens (Aug/2024) -> EXAONE-3.5 7.8B=9T -> EXAONE-3.5 32B=6.5T tokens -> EXAONE 4.0 32B=14T tokens. "We introduce EXAONE 4.5, the first open-weight vision language model developed by LG AI Research. Integrating a dedicated visual encoder... EXAONE 4.5 features 33 billion parameters in total." | 817 | |||||||||||||||||||||||||||||
97 | Muse Spark | Meta AI | https://meta.ai/ | 500 | 50 | MoE | 40,000 | * | 80:1 | 14.9 | 89.5 | 58.4 | synthetic, web-scale | Apr/2026 | 🟢 | D | https://ai.meta.com/blog/introducing-muse-spark-msl/ | Reasoning, SOTA | "This initial model is small and fast by design, yet capable enough to reason through complex questions in science, math, and health." Announce: https://about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs/ "we can reach the same capabilities with over an order of magnitude less compute than our previous model, Llama 4 Maverick. This improvement also makes Muse Spark significantly more efficient than the leading base models available for comparison." | 816 | ||||||||||||||||||||||||||||
98 | Horus 1.0 4B | TokenAI | https://huggingface.co/tokenaii/horus | 4 | Dense | 3,000 | 750:1 | 0.4 | 85 | 60 | 20 | synthetic, web-scale | Apr/2026 | 🟢 | C | https://tokenai.cloud/horus | Reasoning | First open-source AI model from Egypt. | 815 | |||||||||||||||||||||||||||||
99 | Ternary Bonsai 8B | PrismML | https://huggingface.co/collections/prism-ml/ternary-bonsai | 8.19 | Dense | 36,000 | 4,396:1 | 1.8 | synthetic, web-scale | Apr/2026 | 🟢 | A | https://github.com/PrismML-Eng/Bonsai-demo/blob/main/ternary-bonsai-8b-whitepaper.pdf | 1.58-bit ternary {-1,0,+1} weights across embeddings, attention, MLP, and LM head; quantization-aware retrain of Qwen3-8B. 9.4x smaller than FP16; 27 tok/s on iPhone 17 Pro Max. | 814 | |||||||||||||||||||||||||||||||||
100 | Claude Mythos Preview | Anthropic | 10000 | 150 | MoE | 250,001 | * | 26:1 | 166.7 | 94.5 | 64.7 | synthetic, web-scale | Apr/2026 | 🔴 | F | https://www-cdn.anthropic.com/78073f739564e986ff3e28522761a7a0b4484f84.pdf | Reasoning, SOTA | Claude Mythos is suspected to be a Recurrent-Depth Transformer (RDT) — also called a Looped Transformer (LT). https://github.com/kyegomez/OpenMythos "Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available." | 813 |