ABCDEFGHIJKLMNOPQRSTUVWXYZAAABACADAEAFAGAHAIAJAKALAMANAOAPAQARASATAU
1
(910 models)Permalink:
https://lifearchitect.ai/models-table/
Upgrade to Models Table Pro to unlock all columns:
https://lifearchitect.ai/models-table-pro/
The Memo:
https://lifearchitect.ai/memo
Filter using column arrows, or use Menu > Data > Create filter view.
Timeline view:
https://lifearchitect.ai/timeline
ComputeHopperBlackwellRubinMore...
2
ModelLabPlayground
Params
(total, B)
Params
(active, B)
Arch
Tokens
trained (B)
Estimated
total params?
Data
ratio (total)
🔒Training
cost ($)
ALScoreMMLUMMLU
-Pro
GPQAHLE
Training dataset
Announced
Public?🔒License🔒Context
window
Disclosure
score
Paper /
Repo
TagsNotesCount (rough)
🔒Params total
confidence
🔒Params active
confidence
🔒Tokens
confidence
🔒Country🔒Training
hardware
🔒Compute
(FLOPs)
🔒Compute
(Log FLOPs)
🔒Compute
(ZettaFLOPs)
🔒Frontier
compute %
🔒Compute
percentile
🔒H100
hours
🔒H100
energy (MWh)
🔒H100
CO2 Emissions
(tonnes)
🔒B200
hours
🔒B200
cost to train
🔒B200
energy (MWh)
🔒B200
CO2 Emissions
(tonnes)
🔒Rubin
hours (hold)
🔒Rubin
cost to train (hold)
🔒Rubin
energy (MWh)
🔒Rubin
CO2 Emissions
(tonnes)
🔒More private
columns
3
Muse Spark 1.1Meta AIhttps://meta.ai/50050MoE120,000*240:125.862.1
synthetic, web-scale
Jul/2026🟢Dhttps://ai.meta.com/static-resource/muse-spark-1-1-evaluation-reportReasoning, SOTAMultimodal reasoning model from Meta Superintelligence Labs. First model on Meta Model API ($1.25/$4.25 per 1M tokens). Trained for agentic tasks with active 1M-token context management; multi-agent orchestration and thought compression. MSL Research: "Muse Spark 1.1 is a significant upgrade of 1.0, especially on general STEM reasoning, coding, personal and professional agentic tasks. We incorporated more and higher quality data, spent significantly more human research compute and GPU compute with a more stable async RL stack." https://x.com/shuchaobi/status/2075235778561798462910
4
Grok 4.5xAIhttps://grok.com/150075MoE120,00080:144.7
synthetic, web-scale
Jul/2026🟢Chttps://x.ai/news/grok-4-5ReasoningFirst model built jointly with Cursor. V9 foundation architecture; 1.5T params; MoE. Trained on tens of thousands of NVIDIA GB300 GPUs. 80 TPS; $2/$6 per M tokens. Announce: https://cursor.com/blog/grok-4-5909
5
Horus-Hiero-9BTokenAIhttps://huggingface.co/tokenaii/Horus-Hiero-9B9Dense40,0004,445:12.079.378.1
synthetic, web-scale
Jul/2026🟢Chttps://huggingface.co/tokenaii/Horus-Hiero-9BReasoningFine-tune of Qwen3.5-9B for hieroglyph translation and 150+ languages. First hieroglyphic-capable LLM from Egyptian AI startup. Multimodal (text/image/video). 908
6
Hy3Tencenthttps://huggingface.co/tencent/Hy329521MoE40,000136:111.590.453.2
synthetic, web-scale
Jul/2026🟢Chttps://github.com/Tencent-Hunyuan/Hy3ReasoningPost-training upgrade of Hy3 Preview after feedback from 50+ product teams. 192 experts (top-8 activated) with 3.8B MTP layer. Hallucination rate reduced from 12.5% to 5.4%. Scored 2.67/4 vs GLM-5.1 at 2.51/4 in 312 blind expert comparisons. SWE-bench Verified accuracy variance within 4% across scaffoldings (CodeBuddy/Cline/KiloCode). License changed from Community to Apache 2.0.907
7
Leanstral 1.5Mistralhttps://huggingface.co/mistralai/Leanstral-1.5-119B-A6B1196.5MoE15,000127:14.5
synthetic, web-scale
Jul/2026🟢Chttps://mistral.ai/news/leanstral-1-5/ReasoningOpen-source code agent model for Lean 4 theorem proving only (text-only), part of the Mistral Small 4 family (128 experts, 4 active). Saturates miniF2F (100%), solves 587/672 PutnamBench, SOTA on FATE-H (87%) and FATE-X (34%); found 5 previously unknown bugs across 57 real repositories.906
8
Hierarchos 232MIndependenthttps://github.com/necat101/Hierarchos0.232Dense0.131:10.001web-scaleJul/2026🟢Chttps://github.com/necat101/Hierarchos/blob/main/HIERARCHOS_FINDINGS_PAPER.mdExperimental hybrid non-Transformer (RWKV v8 + Titans memory + HRM). 13 epochs on Alpaca SFT on RTX 6000 Blackwell. 'Dataset: netcat420/Experiment_0.1 (Alpaca format)' + 'Training: 13 epochs'. Standard Alpaca ~52K samples × ~200 tokens/sample × 13 epochs ≈ 135M. Smoke-test (n=100): ARC Easy 0.36, HellaSwag 0.37, TruthfulQA 0.22. Announce: https://www.reddit.com/r/MachineLearning/comments/1um123n/hierarchos_preliminary_findings_from_a_232m/905
9
Laguna XS 2.1Poolsidehttps://huggingface.co/poolside/Laguna-XS-2.133.43MoE30,000899:13.353
synthetic, web-scale
Jul/2026🟢Ahttps://poolside.ai/assets/laguna/laguna-m1-xs2-technical-report.pdfReasoning33B total / 3B active MoE for agentic coding on a local machine; upgrade of XS.2 with +5.4pt SWE-bench Multilingual. SWE-bench Verified 70.9, SWE-bench Multilingual 63.1, SWE-Bench Pro 47.6, Terminal-Bench 2.0 37.5. MMLU-Pro from XS.2 base model report. Announce: https://poolside.ai/blog/introducing-laguna-xs-2-1904
10
Nemotron-Labs-TwoTower-30B-A3B
NVIDIAhttps://huggingface.co/nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16606MoE27,100452:14.378.2460.93
synthetic, web-scale
Jul/2026🟢Ahttps://arxiv.org/abs/2606.26493DiffusionTwo-tower block-wise autoregressive diffusion LLM built on Nemotron-3-Nano-30B-A3B. Frozen AR context tower + trainable diffusion denoiser tower. Retains 98.7% of AR baseline quality at 2.42x generation throughput. 128K context. NVIDIA Nemotron Open Model License.903
11
TabFM 1.0Googlehttps://huggingface.co/google/tabfm-1.0.0-pytorch1.6Dense600*375:10.1syntheticJun/2026🟢Dhttps://github.com/google-research/tabfmZero-shot tabular foundation model for classification and regression via in-context learning. Trained entirely on hundreds of millions of synthetic datasets. Blog: "trained entirely on hundreds of millions of synthetic datasets." Comparable models: TabPFN v2 trained on 130M datasets; TabICLv2 on ~33M+ datasets across 550K steps. "Hundreds of millions" → ~300M datasets. TabICLv2's curriculum uses datasets from 1K to 60K rows; average ~2K. So ~300M × 2K rows = ~600B tabular data points. This parallels TimesFM's "100B time-points" convention for non-text foundation models. #1 on TabArena. Being integrated into Google BigQuery.902
12
Claude Sonnet 5Anthropichttps://claude.ai/ 100020MoE80,000*80:129.88957.4
synthetic, web-scale
Jun/2026🟢Dhttps://www-cdn.anthropic.com/d9bb04416ffe1352af84721476c1fa9994c07fde/Claude%20Sonnet%205%20System%20Card.pdfReasoning1M context. Announce: https://www.anthropic.com/news/claude-sonnet-5. Showing GMMLU (Global MMLU by Cohere).901
13
openPangu-2.0-FlashHuaweihttps://huggingface.co/openpangu/openPangu-2.0-Flash926MoE34,000370:15.983.7
synthetic, web-scale
Jun/2026🟢Ahttps://ai.gitcode.com/ascend-tribe/openPangu-2.0-FlashReasoningFirst frontier-scale MoE model trained entirely on non-NVIDIA hardware (Ascend 910B NPUs). 512K context. MLA + DSA/SWA hybrid attention (1:2 ratio). 3-head MTP (Multi-Token Prediction). Muon optimizer. Post-training: SFT + multi-objective RL + online policy distillation (OPD). GPQA-Diamond=83.7 (Thinking; Avg@4). AIME 2026=93.3 (Avg@16). SWE-bench Verified=63.1 (Avg@3). License: openPangu Model License Agreement v2.0.900
14
LongCat-2.0Meituanhttps://longcat.ai160048MoE35,00022:124.988.9
synthetic, web-scale
Jun/2026🟢Ahttps://longcat.chat/blog/longcat-2.0/ReasoningTrained entirely on domestic AI ASIC superpods with no NVIDIA GPUs. LongCat Sparse Attention for 1M context. MOPD post-training from agent, reasoning, and interaction expert groups.899
15
Agents-A1Shanghai AI Labhttps://huggingface.co/InternScience/Agents-A1353MoE40,0001,143:13.947.6
synthetic, web-scale
Jun/2026🟢Chttps://arxiv.org/abs/2606.30616Reasoning35B MoE agentic model fine-tuned from Qwen3.5-35B-A3B via three-stage recipe: full-domain SFT, domain-level teacher training, and multi-teacher on-policy distillation. Matches or outperforms 1T-parameter models on long-horizon agent benchmarks.898
16
GPT-5.6 SolOpenAIhttps://chatgpt.com/3000150MoE200,000*67:181.6
synthetic, web-scale
Jun/2026🟢Dhttps://deploymentsafety.openai.com/gpt-5-6-previewReasoning, SOTAFlagship of GPT-5.6 family (Sol/Terra/Luna). SOTA Terminal-Bench 2.1. CTF saturated at 96.7%. $5/$30 per 1M tokens. Limited preview to trusted partners; broad release planned. Developer log analysis and pre-release reports suggest a 1.5M token context window, not yet officially confirmed by OpenAI in the blog or system card. Max and Ultra reasoning modes. 'We will share an expanded suite of evaluation results when we make the model broadly available.'897
17
Ornith-1.0-397BDeepReinforcehttps://huggingface.co/collections/deepreinforce-ai/ornith-1039717MoE36,00091:112.6
synthetic, web-scale
Jun/2026🟢Ahttps://deep-reinforce.com/ornith_1_0.htmlReasoningSelf-improving RL framework for agentic coding. Post-trained on Qwen 3.5-397B-A17B. 82.4 SWE-Bench Verified, 77.5 Terminal-Bench 2.1.896
18
Unlimited-OCRBaiduhttps://huggingface.co/baidu/Unlimited-OCR30.5MoE6,0002,000:10.4
synthetic, web-scale
Jun/2026🟢Chttps://arxiv.org/abs/2606.23050End-to-end OCR model building on DeepSeek-OCR; replaces all decoder attention with Reference Sliding Window Attention (R-SWA) for constant KV cache; 3B total / 500M active MoE; scores 93.23 on OmniDocBench v1.5 (+6% over DeepSeek-OCR); parses dozens of pages in a single forward pass at 32K max length.895
19
Seed2.1 ProByteDancehttps://exp.volcengine.com/ark?csid=excs-202602141507-%5BJk-cTyC-D-fHC56fiUG_K%5D&mode=chat&modelId=doubao-seed-2-1-pro-26062850050MoE30,000*60:112.955.7
synthetic, web-scale
Jun/2026🟢Dhttps://lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlaukjlkulzlp/seed2.1/Seed2_1_Model_Card.pdfReasoningAgentic productivity model. Ranked 8th on Code Arena: Frontend (1539). Compared against Claude Opus 4.7 and GPT-5.5.894
20
QUEST-35B-RLOSU NLPhttps://huggingface.co/osunlp/QUEST-35B-RL353MoE40,0001,143:13.937.9
synthetic, web-scale
Jun/2026🟢Chttps://arxiv.org/abs/2605.24218Qwen3.5-35B-A3B base, Open deep research agent family (2B–35B) trained with fully synthetic rubric-tree tasks on 32 H100s. Approaches or surpasses frontier closed-source agents across eight deep research benchmarks. Announce: https://x.com/ysu_nlp/status/2067380438134624742 893
21
Glimmer-1-BaseGlint Researchhttps://huggingface.co/Glint-Research/Glimmer-1-Base0.0000119Dense0.000543:10.000web-scaleJun/2026🟢Ahttps://huggingface.co/Glint-Research/Glimmer-1-Base11.9K-parameter (0.0000119B) experimental micro-model trained on 500K tokens of FineWeb-Edu. Llama-style transformer exploring the lower bound of useful language model scale. Base only, no SFT. Trained on FineWeb-Edu on a single RTX 4070 SUPER.892
22
GLM-5.2Z.AIhttps://huggingface.co/zai-org/GLM-5.274440MoE28,50039:115.391.254.7
synthetic, web-scale
Jun/2026🟢Ahttps://arxiv.org/abs/2602.15763Reasoning1M-token context (up from 200K in GLM-5.1); 131K max output. Trained entirely on Huawei Ascend 910B; no NVIDIA hardware. Two thinking modes (High; Max).891
23
VibeThinker-3BWeiboAI (Sina Weibo)https://huggingface.co/WeiboAI/VibeThinker-3B3Dense5,5001,834:10.470.2
synthetic, web-scale
Jun/2026🟢Chttps://arxiv.org/abs/2606.16140Reasoning3B dense reasoning model scoring 94.3 on AIME26 and 80.2 on LiveCodeBench v6; matches 100x+ larger models on verifiable reasoning tasks. Built on Qwen2.5-Coder-3B via curriculum SFT + multi-domain RL + offline self-distillation.890
24
SubQ 1.1 SmallSubquadratichttps://subq.ai/request-early-access70Dense13,000*186:13.285.4
synthetic, web-scale
Jun/2026🟢Dhttps://subq.ai/docs/subq-1-1-small-model-card.pdfReasoningSSA (Subquadratic Sparse Attention). Near-perfect NIAH retrieval to 12M tokens. GPQA Diamond 85.4%. LiveCodeBench v6 pass@4=89.7%. RULER 128K=99.12%. AutomationBench Finance=13%. Third-party verified by Appen. 64.5x less compute than dense attention at 1M tokens.889
25
Rio-3.5-Open-397BIplanRIOhttps://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B39717MoE36,00091:112.68890.936.5
synthetic, web-scale
Jun/2026🟢Chttps://huggingface.co/prefeitura-rio/Rio-3.5-Open-397BReasoningMerge of Qwen 3.5 397B + Next-N2-Pro. IplanRIO is Rio de Janeiro municipal IT company. Features SwiReasoning: dynamic latent/explicit reasoning via entropy-based confidence signals. SWE-Bench Verified=80.2. https://github.com/nex-agi/Nex-N2/issues/4888
26
openPangu-2.0-ProHuaweipending 30/jun50518MoE19,00038:110.3
synthetic, web-scale
Jun/2026🟢Chttps://gitcode.com/ascend-tribeOpen-source MoE with record 28:1 sparsity ratio. DSA+SWA hybrid attention architecture. Optimized for Ascend NPU; 2x single-card throughput vs mainstream open-source models. Open-sourcing 7 components from 30/Jun/2026. Dataset: Predecessor openPangu-Ultra-MoE-718B (718B/39B active) trained on ~19T tokens.887
27
Kimi-K2.7-CodeMoonshot AIhttps://huggingface.co/moonshotai/Kimi-K2.7-Code100032MoE30,50031:118.4
synthetic, web-scale
Jun/2026🟢Ahttps://huggingface.co/moonshotai/Kimi-K2.7-CodeReasoningCoding-focused agentic model built upon Kimi K2.6. Reduces thinking-token usage by ~30% compared to K2.6. 15.5T is verified base pretraining only; K2.5→K2.6→K2.7 continued training adds undisclosed tokens886
28
Nex-N2-ProNex AGIhttps://huggingface.co/nex-agi/Nex-N2-Pro39717MoE36,00091:112.690.7
synthetic, web-scale
Jun/2026🟢Chttps://github.com/nex-agi/Nex-N2ReasoningPost-trained on Qwen3.5-397B-A17B. "An agentic model with Agentic Thinking." GPQA Diamond up from 88.4 (base Qwen3.5) to 90.7 (+2.3 from post-training). SWE-Bench Verified=80.8, Terminal-Bench 2.1=75.3, SWE-Bench Pro=58.8. Competitive with GPT-5.5 and Opus 4.7 on coding and agentic benchmarks.885
29
DiffusionGemma 26B A4B IT
Google DeepMindhttps://huggingface.co/google/diffusiongemma-26B-A4B-it25.23.8MoE14,000556:12.077.673.211.9web-scaleJun/2026🟢Chttps://huggingface.co/google/diffusiongemma-26B-A4B-it
Reasoning, Diffusion
Experimental discrete diffusion model on Gemma 4 26B A4B MoE backbone. Generates 256-token blocks in parallel via iterative denoising (1000+ tok/s on H100, 700+ on RTX 5090). Bidirectional attention enables self-correction. Quality lower than standard Gemma 4; designed for speed-critical local workflows.884
30
Apodex-1.0-HApodex AIhttps://apodex.ai39717MoE36,000*91:112.660.8
synthetic, web-scale
Jun/2026🟢Dhttps://www.apodex.com/blog/apodex-1.0ReasoningVerification-centric deep-research agent team on Qwen3.5 base. Heavy-duty mode coordinates up to 150 sub-agents over 15,000 steps. SOTA on BrowseComp (90.3), DeepSearchQA (94.4), FrontierScience-Research (46.7).883
31
Claude Fable 5Anthropichttps://claude.ai/10000150MoE250,000*25:1166.794.164.5
synthetic, web-scale
Jun/2026🟢Dhttps://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdfReasoning, SOTAMythos-class model made safe for general use. Same underlying model as Claude Mythos 5 with safety classifiers (fallback to Opus 4.8 in <5% of sessions for cyber, bio/chem, distillation). Pricing $10/$50 per Mtok. API: claude-fable-5.882
32
North-Mini-Code-1.0Coherehttps://huggingface.co/CohereLabs/North-Mini-Code-1.0303MoE12,000400:12.0
synthetic, web-scale
Jun/2026🟢Chttps://huggingface.co/blog/CohereLabs/introducing-north-mini-code30B-A3B MoE (128 experts, 8 active per token) optimized for agentic software engineering. First model in Cohere's North family. Artificial Analysis Coding Index: 33.4. 256K context, 64K output.881
33
AFM 3 Core AdvancedApplehttps://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models204MoE25,0001,250:12.4
synthetic, web-scale
Jun/2026🟢Chttps://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-modelsMost powerful Apple on-device model. 20B params stored in flash (NAND); 1–4B activated per prompt via Instruction-Following Pruning (IFP). Natively multimodal (text, image, audio). Built with Google on cloud TPUs. 2026 blog states 'we significantly scaled pre-training on the latest generation of cloud TPU accelerators' and 'all models shared a common initial foundation.' Conservative 15T estimate accounts for scaling over 14T+ base + multimodal tokens.880
34
AFM 3 Cloud ProApplehttps://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models120060MoE60,000*50:128.3
synthetic, web-scale
Jun/2026🟢Dhttps://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-modelsApple–Google–NVIDIA collaboration. Based on custom 1.2T-parameter Gemini model with Apple's own pre-training and post-training. Runs on NVIDIA GPUs in Google Cloud via extended Private Cloud Compute. 'Our most capable server-based model, which powers our most demanding use cases, like agentic tool use and complex reasoning.' Tech report planned for summer 2026. Bloomberg, Mark Gurman, Nov 5 2025: "Apple Inc. is planning to pay about $1 billion a year for an ultrapowerful 1.2 trillion parameter artificial intelligence model developed by Alphabet Inc.'s Google" URI: https://www.bloomberg.com/news/articles/2025-11-05/apple-plans-to-use-1-2-trillion-parameter-google-gemini-model-to-power-new-siri "Based on Gemini foundation… Apple did their own pre-training, post-training" — Max Weinbach tweet, Jun 8 2026, quoted in wccftech: "Apple just clarified AFM Cloud is Apple's own model, trained with Gemini outputs / AFM local models are entirely Apple models / AFM Cloud Pro seems to be based on Gemini foundation and data, but Apple did their own pre-training, post-training, RL, etc" URI: https://wccftech.com/apple-removes-the-fog-around-its-new-cloud-based-and-20-billion-parameter-on-device-ai-models-brushes-aside-googles-contributions-while-hyping-nvidias/.879
35
Macaron-V1-Preview-749B
Mind Labhttps://macaron-model-previews.macaron.im/74941MoE28,50039:115.4
synthetic, web-scale
Jun/2026🟢Ahttps://macaron.im/mindlab/research/macaron-v1-preview749B Mixture-of-LoRA agent model post-trained from GLM-5.1 (744B frozen base + 5 × 1B specialist LoRAs for chat, personal-life, coding, Generative UI, and OpenClaw tasks). Router Tool routes between adapters. SWE-bench Verified=78.1.878
36
Gemma 4 12BGoogle DeepMindhttps://huggingface.co/google/gemma-4-12B-it12Dense14,0001,167:11.477.278.85.2web-scaleJun/2026🟢Chttps://arxiv.org/abs/2607.02770ReasoningEncoder-free multimodal (text, image, audio) dense model with configurable thinking mode and 256K context.877
37
Aion-1.0-PlanMicrosoft14Dense14,0001,000:11.5
synthetic, web-scale
Jun/2026🟢Dhttps://blogs.windows.com/windowsdeveloper/2026/06/02/build-2026-furthering-windows-as-the-trusted-platform-for-development/ReasoningOn-device reasoning and tool-calling SLM that ships in-box as part of Windows on capable devices, enabling fully local agentic workflows. "Enables applications to reason over user intent, invoke tools, manage files and orchestrate sub-agents, bringing fully agentic workflows onto the device." Announced at Build 2026; available in the coming months.876
38
Aion-1.0-InstructMicrosofthttps://microsoftedge.github.io/Demos/built-in-ai/playgrounds/prompt-api/2Dense8,000*4,000:10.4
synthetic, web-scale
Jun/2026🟢Dhttps://blogs.windows.com/msedgedev/2026/06/02/expanding-on-device-ai-in-microsoft-edge-new-models-and-apis-for-the-web/Pre-release small language model for on-device AI in Microsoft Edge (Canary/Dev), powering the Prompt and Writing Assistance APIs. Successor to Phi-4-mini (4B); "smaller, faster, and more efficient," supports CPU inference for devices without a GPU. Planned open-source release on Hugging Face in July 2026.875
39
MAI-Code-1-FlashMicrosofthttps://github.blog/changelog/2026-06-02-mai-code-1-flash-is-now-available-for-github-copilot/30Dense15,000*500:12.2web-scaleJun/2026🟢Dhttps://microsoft.ai/news/introducingmai-code-1-flash/ReasoningLightweight agentic coding model from Microsoft AI, built end-to-end on clean and appropriately licensed data, trained directly with GitHub Copilot harnesses. Adaptive solution-length control: solves harder problems with up to 60% fewer tokens. Outperforms Claude Haiku 4.5 across SWE-Bench Verified, SWE-Bench Pro (51.2% vs 35.2%), SWE-Bench Multilingual, and Terminal Bench 2. Available in VS Code GitHub Copilot.874
40
MAI-Thinking-1Microsofthttps://microsoft.ai/news/introducing-mai-thinking-1/100035MoE33,50034:119.38584.2web-scaleJun/2026🟢Ahttps://microsoft.ai/wp-content/uploads/2026/06/main_20260602_2.pdfReasoningMicrosoft AI's reasoning model. 35B-active, ~1T-total parameters sparse MoE. Trained from the ground up without distillation from third-party models, on clean and commercially licensed data. Matches Claude Opus 4.6 on SWE-Bench Pro and preferred over Claude Sonnet 4.6 in blind human side-by-side evaluations. AIME 2025=97.0, AIME 2026=94.5.873
41
KeyLM-75M-InstructIndependenthttps://huggingface.co/Eclipse-Senpai/KeyLM-75M-Instruct0.0753Dense18240:10.00424web-scaleJun/2026🟢Ahttps://huggingface.co/Eclipse-Senpai/KeyLM-75M-Instruct75M-param from-scratch small LM; competitive on IFEval vs SmolLM-135M-Instruct at half the size. "trained completely on kaggle (tpu v5e-8)" Announce: https://www.reddit.com/r/LocalLLaMA/comments/1tuyb8s/i_trained_a_75m_parameter_llm_from_scratch_on_18b/872
42
Cosmos 3 SuperNVIDIAhttps://huggingface.co/nvidia/Cosmos3-Super6432MoE2004:10.4specialJun/2026🟢Chttps://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdfSOTAOmnimodal world model for Physical AI; dual-tower mixture-of-transformers (reasoner + generator) initialized from Qwen3-VL-32B. Dataset: ‘two epochs over the full pre-training mixture’ with sequences ‘at most 16k tokens.’ Conservative avg ~4K tokens/sample × 22M × 2 epochs ≈ 176B pretrain + ~9B SFT ≈ ~185B; rounded to 200B. Excludes generator-pathway vision/audio/action tokens (hundreds of millions of images and videos, not directly token-comparable to LLM corpora) and the prior Qwen3-VL-32B initialization tokens (not disclosed).871
43
Mellum2-12B-A2.5B-Thinking
JetBrainshttps://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Thinking122.5MoE10,600884:11.257.6
synthetic, web-scale
Jun/2026🟢Ahttps://arxiv.org/abs/2605.31268ReasoningOpen-weight 12B MoE (64 experts, 8 active) language model specialised in software engineering; successor to the 4B dense Mellum.870
44
Qwen3.7-PlusAlibabahttps://chat.qwen.ai/48035MoE40,000*84:114.688.590.334.7
synthetic, web-scale
Jun/2026🟢Dhttps://qwen.ai/blog?id=qwen3.7-plusReasoningMultimodal agent model unifying vision and language; operates GUI and CLI within a single agent loop869
45
Nemotron 3 UltraNVIDIAhttps://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16#nvidia-nemotron-3-ultra-550b-a55b-bf1655055MoE25,00046:112.489.186.88737.4
synthetic, web-scale
Jun/2026🟢Chttps://arxiv.org/abs/2512.20856NVIDIA’s largest open model: 550B total parameters with up to 55B active per token via a hybrid Mamba-Transformer MoE architecture. Most intelligent US open weights model per Artificial Analysis (Intelligence Index 48).868
46
MiniMax-M3MiniMaxhttps://huggingface.co/MiniMaxAI/MiniMax-M342823MoE100,000234:121.8
synthetic, web-scale
Jun/2026🟢Chttps://www.minimax.io/blog/minimax-m3Reasoning, SOTA"M3 is a model that has undergone mixed-modality training from Step 0... After rebuilding the entire data pipeline for this data, we are now able to scale the training data to the order of 100 trillion tokens." First open-weight model combining frontier coding, 1M-token MSA context, and native multimodality. SWE-Bench Pro: 59.0; BrowseComp: 83.5 (surpasses Opus 4.7 at 79.3). Tech report and weights to be released within 10 days of launch.867
47
Step 3.7 FlashStepFunhttps://huggingface.co/stepfun-ai/Step-3.7-Flash19811MoE24,000122:17.349.7
synthetic, web-scale
May/2026🟢Ahttps://github.com/stepfun-ai/Step-3.7-FlashReasoningA high-efficiency Flash model for real-world agents.866
48
LFM2.5-8B-A1BLiquid AIhttps://huggingface.co/LiquidAI/LFM2.5-8B-A1B8.31.5MoE38,0004,579:11.9
synthetic, web-scale
May/2026🟢Ahttps://www.liquid.ai/blog/lfm2-5-8b-a1bReasoningEdge MoE for fast on-device tool calling; 128K context, reasoning-only model. Highlights: IFEval 91.84, MATH500 88.76, BFCLv3 64.79, Tau2-Telecom 88.07.865
49
Claude Opus 4.8Anthropichttps://claude.ai/5000150MoE80,000*16:166.793.657.9
synthetic, web-scale
May/2026🟢Dhttps://www.anthropic.com/claude-opus-4-8-system-cardReasoning, SOTAAnnounce: https://www.anthropic.com/news/claude-opus-4-8 HLE=with tools (49.8 no tools). Same price as Opus 4.7 ($5/$25). Params/tokens carried from Opus-class estimate (Opus 4.7).864
50
ESMC 6BBiohubhttps://huggingface.co/biohub/ESMC-6B6Dense6,6001,100:10.7specialMay/2026🟢Ahttps://biohub.ai/papers/esm_protein.pdf“Language Modeling Materializes a World Model of Protein Biology”. Protein language model, 80 layers, 2.37e23 training FLOPs. Trained on ~2.8B protein sequences (UniRef + MGnify + JGI, clustered at 70% identity). Tokens back-calculated from disclosed compute: training FLOPs of 2.37e23 reported on the HF model card, divided by 6N (Kaplan/Chinchilla rule for transformer training), gives 2.37e23 ÷ (6 × 6e9) ≈ 6.6T tokens. Released alongside ESMFold2 and ESM Atlas (6.8B sequences / 1.1B predicted structures). MIT license.863
51
MiniCPM5-1BOpenBMBhttps://huggingface.co/spaces/openbmb/MiniCPM5-1B-Demo1.08Dense8,0007,408:10.348.8526.26
synthetic, web-scale
May/2026🟢Ahttps://huggingface.co/openbmb/MiniCPM5-1BReasoningthe first model in the MiniCPM5 series. It is a dense 1B Transformer built for on-device, local deployment, and resource-constrained scenarios, reaching 1B-class open-source SOTA. Hybrid reasoning with <think> template. 1,080,632,832 params, 24 layers, GQA 16Q/2KV, 131K context. Post-training: 200B deep-thinking SFT + 200B hybrid-thinking SFT + RL + OPD. Pretraining token count not disclosed; estimate aligned with MiniCPM4 paper's 8T-token series budget (UltraClean/Ultra-FineWeb data-efficient pipeline). MMLU-Redux=70.06, MATH-500=91.6, AIME-2025=40.42, BBH=71.89, IFEval=80.41 (Thinking mode).862
52
Gated DeltaNet-2NVIDIAhttps://github.com/NVlabs/GatedDeltaNet-21.3Dense10077:10.04web-scaleMay/2026🟢Ahttps://github.com/NVlabs/GatedDeltaNet-2/blob/main/paper/GDN2_paper.pdfLinear-attention architecture decoupling channel-wise erase and write gates; 1.3B params trained on 100B tokens of FineWeb-Edu; outperforms Mamba-2, Gated DeltaNet, KDA, and Mamba-3 on language modeling, RULER, and retrieval.861
53
Command A+Coherehttps://huggingface.co/CohereLabs/command-a-plus-05-2026-bf1621825MoE20,00092:17.0
synthetic, web-scale
May/2026🟢Chttps://huggingface.co/CohereLabs/command-a-plus-05-2026-bf16"open source model with 25 billion active parameters and 218B total parameters model optimized for agentic, multilingual, and reasoning-heavy tasks with a focus on enterprise performance, while also providing support for vision inputs". 128 experts, 8 active per token + 1 shared expert. 128K context. 48 languages. Announce: https://cohere.com/blog/command-a-plus860
54
Qwen3.7-MaxAlibabahttps://chat.qwen.ai/2000100MoE40,000*20:129.889.692.453.5
synthetic, web-scale
May/2026🟢Dhttps://qwen.ai/blog?id=qwen3.7Reasoning"Qwen3.7-Max, our latest proprietary model designed for the agent era." 35-hour autonomous kernel optimization run with 1,000+ tool calls; 10.0x geomean speedup over Triton reference. Available soon via Alibaba Cloud Model Studio.859
55
HRM-Text-1BSapient Intelligencehttps://huggingface.co/sapientinc/HRM-Text-1B1Dense160160:10.0460.7
synthetic, web-scale
May/2026🟢Ahttps://github.com/sapientinc/HRM-TextReasoning"1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning." Pretraining cost ~$1472 on 16 H100s. 40B unique tokens × 4 epochs = 160B total.858
56
Nemotron-Labs-Diffusion-14B
NVIDIAhttps://huggingface.co/nvidia/Nemotron-Labs-Diffusion-14B14Dense4,345*311:10.882.5154.55
synthetic, web-scale
May/2026🟢Dhttps://d1qx31qr3h6wln.cloudfront.net/publications/Nemotron_Diffusion_Tech_Report_v1.pdfDiffusionTri-mode LM unifying AR, diffusion, and self-speculation decoding within a single architecture. Explicit Nemotron training is disclosed at 1T (Stage 1 AR) + 300B (Stage 2 joint) + 45B (SFT) = 1,345B. Base init from Ministral3-14B (Liu et al., arXiv 2601.08584, 2026)≈3T (disclosed).857
57
Gemini 3.5 FlashGoogle DeepMindhttps://gemini.google.com/50025MoE100,000200:123.640.2
synthetic, web-scale
May/2026🟢Ahttps://deepmind.google/models/model-cards/gemini-3-5-flash/“Gemini 3.5 Flash delivers intelligence that rivals large flagship models on multiple dimensions, at the speeds you have come to expect from the Flash series.” Beats Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), MCP Atlas (83.6%), CharXiv Reasoning (84.2%). Pro coming next month. Announce: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/856
58
ZAYA1-8B-Diffusion-Preview
Zyphrahttps://huggingface.co/Zyphra/ZAYA1-8B8.40.76MoE15,1001,798:11.294
synthetic, web-scale
May/2026🟢Ahttps://www.zyphra.com/post/zaya1-8b-diffusion-previewDiffusionFirst MoE diffusion model converted from an autoregressive LLM, and first diffusion-LM trained on AMD. Built from ZAYA1-8B base via TiDAR-style conversion: 600B tokens at 32k + 500B context extension to 128k + diffusion SFT. 4.6x decoding speedup (lossless sampler), 7.7x (mixed-logits sampler). Diffuses blocks of 16 tokens simultaneously.855
59
Intern-S2-PreviewShanghai AI Laboratory/SenseTimehttps://huggingface.co/internlm/Intern-S2-Preview353MoE41,0001,172:14.08818.07
synthetic, web-scale
May/2026🟢Ahttps://huggingface.co/internlm/Intern-S2-PreviewReasoning"an efficient 35B scientific multimodal foundation model... continued pretrained from Qwen3.5." 35B-A3B MoE. Tokens estimate: Qwen3.5 base ~36T + ~5T scientific continued pretraining (matching Intern-S1-Pro pattern) ≈ 41T. Announce: https://github.com/InternLM/Intern-S1854
60
Ring-2.6-1TInclusion AIhttps://huggingface.co/inclusionAI/Ring-2.6-1T100050MoE20,75021:115.288.27
synthetic, web-scale
May/2026🟢Chttps://arxiv.org/abs/2606.15079ReasoningReasoning sibling of Ling-2.6-1T. Async RL training + IcePop algorithm. Two reasoning effort levels: high and xhigh. xhigh: AIME 26=95.83, ARC-AGI-V2=66.18, GPQA Diamond=88.27. high: PinchBench=87.60, ClawEval=63.82, Tau2-Bench Telecom=95.32.853
61
TML-Interaction-SmallThinking Machines Lab27612MoE30,000109:19.6
web-scale, audio, video
May/2026🟡Dhttps://thinkingmachines.ai/blog/interaction-models/Native multimodal (audio, video, text) interaction model with time-aligned 200ms micro-turns. “TML-Interaction-Small dominates interaction quality while being more intelligent than any non thinking model.” Research preview; larger models planned later in 2026.852
62
NeedleCactus-Computehttps://huggingface.co/Cactus-Compute/needle0.026Dense2027,770:10.008
synthetic, web-scale
May/2026🟢Ahttps://github.com/cactus-compute/needleDistilled from Gemini 3.1 into a 26M-parameter "Simple Attention Network" (attention + gating, no MLPs/FFNs). 6000 tok/s prefill, 1200 tok/s decode on consumer devices. Pretrained on 16 TPU v6e for 200B tokens (27 hrs), post-trained on 2B tokens of single-shot function-calling data (45 min) synthesized via Gemini across 15 tool categories. MIT licensed. Weights: https://huggingface.co/Cactus-Compute/needle851
63
NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B
NVIDIAhttps://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B-BF16303.6MoE25,160839:12.978.6372.1
synthetic, web-scale
May/2026🟢Ahttps://arxiv.org/abs/2511.16664"NVIDIA-Nemotron-Labs-3-Elastic-30B-A3B-BF16 is a 3-in-1 elastic large language model (LLM) developed by NVIDIA. It contains three nested model variants (30B, 23B, and 12B parameters) within a single BF16 checkpoint, all sharing the same parameter space."850
64
ZAYA1-74B-PreviewZyphrahttps://huggingface.co/Zyphra/ZAYA1-74B-preview744MoE18,000244:13.884.483.8
synthetic, web-scale
May/2026🟢Ahttps://www.zyphra.com/zaya1-8b-technical-reportReasoning"Pre-RL reasoning base checkpoint (no instruction or RL post-training). Trained end-to-end on AMD MI300x hardware. Uses CCA attention with sliding window attention hybrid and 256k context. "ZAYA1-74B-Preview is a pre-RL reasoning-base checkpoint, released under an Apache 2.0 license."849
65
SubQ 1M-PreviewSubquadratichttps://subq.ai/request-early-access70Dense12,000*172:13.1
synthetic, web-scale
May/2026🟢Dhttps://subq.ai/how-ssa-makes-long-context-practicalFirst fully subquadratic LLM. SSA (Subquadratic Sparse Attention). 12M token research context, 1M production. SWE-Bench Verified=81.8.848
66
Llama 3.1 8B + CUA (quantum)
Multiverse Computing8Dense15,0001,875:11.2
synthetic, web-scale
May/2026🟡Bhttps://arxiv.org/abs/2605.05914QuantumFirst end-to-end quantum enhancement of a production-scale LLM on real superconducting hardware (156-qubit IBM Quantum System Two); Cayley unitary adapters add 6,000 params and reduce Llama 3.1 8B perplexity by 1.4%.847
67
ZAYA1-8BZyphrahttps://huggingface.co/Zyphra/ZAYA1-8B8.40.76MoE14,0001,667:11.174.271
synthetic, web-scale
May/2026🟢Ahttps://www.zyphra.com/zaya1-8b-technical-reportReasoningFirst MoE model pretrained, midtrained, and SFT’d entirely on AMD Instinct MI300; 760M active params; Apache-2.0. Announce: https://www.zyphra.com/post/zaya1-8b846
68
GENE-26.5Genesis AIhttps://www.genesis.ai/blog/gene-26-5-advancing-robotic-manipulation-to-human-level30Dense10,000*334:11.8roboticsMay/2026🟡Dhttps://www.genesis.ai/blog/gene-26-5-advancing-robotic-manipulation-to-human-level"our first robotic foundation model system and the initial public release in the GENE family… designed to push general-purpose robotic manipulation towards human-level capability." 200,000+ hours of multimodal data (vision, hand state, language, tactile, robot controls). 200,000 hours × ~2M tokens/hour ≈ 400B multimodal tokens, + internet language/video pretraining priors ≈ 10T total.845
69
GPT-5.5 InstantOpenAIhttps://chatgpt.com/30015MoE114,000*380:119.585.6
synthetic, web-scale
May/2026🟢Dhttps://openai.com/index/gpt-5-5-instant/Reasoning, SOTA"smarter and more accurate, with clearer, more concise answers that feel better tailored to you. Because Instant is the daily driver for hundreds of millions of people..."844
70
MAMMALIBMhttps://huggingface.co/ibm/biomed.omics.bl.sm.ma-ted-458m0.458Dense7001,529:12.3specialMay/2026🟢Ahttps://www.nature.com/articles/s44386-026-00047-4"MAMMAL (Molecular Aligned Multi Modal Architecture and Language), a foundation model for cross-modal learning, designed to address the challenges associated with drug discovery tasks." 2 Billion total samples × ~350 average tokens/sample ≈ 700 Billion tokens.843
71
ERNIE-5.1-PreviewBaiduhttps://ernie.baidu.com/80060MoE100,000125:129.884.391
synthetic, web-scale
Apr/2026🟢Chttps://ernie.baidu.com/blog/posts/ernie-5.1-preview-0430-release-on-lmarena/Reasoning"ERNIE-5.1-Preview builds on the pre-training foundation of ERNIE-5.0 while compressing total parameters to approximately 1/3 and active parameters to approximately 1/2, achieving leading performance at its model scale using only about 6% of the pre-training cost of comparable models."842
72
Granite-4.1-30BIBMhttps://huggingface.co/ibm-granite/granite-4.1-30b30Dense15,000500:12.380.1664.0945.76
synthetic, web-scale
Apr/2026🟢Ahttps://huggingface.co/blog/ibm-granite/granite-4-1Reasoning"Granite 4.1 is a family of dense, decoder‑only LLMs (3B, 8B, and 30B) trained on ~15T tokens using a multi‑stage pre‑training pipeline, including long‑context extension of up to 512K tokens."841
73
Mistral Medium 3.5Mistralhttps://chat.mistral.ai/chat128Dense12,00094:14.1
synthetic, web-scale
Apr/2026🟢Chttps://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5"a dense 128B model with a 256k context window, handling instruction-following, reasoning, and coding in a single set of weights"840
74
Nemotron 3 Nano Omni
NVIDIAhttps://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16303MoE25,700857:12.977.372.2
synthetic, web-scale
Apr/2026🟢Ahttps://arxiv.org/abs/2604.24954ReasoningBuilt on the highly efficient Nemotron 3 Nano 30B-A3B backbone, Nemotron 3 Nano Omni natively supports audio inputs alongside text, images, and video. ~717B tokens (multimodal post-training) on top of ~25T base LLM pretraining, ~25.7T total.839
75
Laguna XS.2Poolsidehttps://huggingface.co/poolside/Laguna-XS.2333MoE10,000304:11.9
synthetic, web-scale
Apr/2026🟢Chttps://poolside.ai/blog/introducing-laguna-xs2-m1"Laguna XS.2, as open-weights. 33B total parameters, 3B active. Apache 2.0."838
76
Laguna M.1Poolsidehttps://platform.poolside.ai/22523MoE10,00045:15.0
synthetic, web-scale
Apr/2026🟢Chttps://poolside.ai/blog/introducing-laguna-xs2-m1"Laguna M.1 is a 225B total parameter model with 23B activated parameters, built for agentic coding and long-horizon work."837
77
DeepSeek-V4-ProDeepSeek-AIhttps://huggingface.co/deepseek-ai/DeepSeek-V4-Pro160049MoE33,00021:124.290.187.590.137.7
synthetic, web-scale
Apr/2026🟢Ahttps://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdfSOTA, ReasoningMoE with 1.6T total / 49B active parameters, 1M-token context, trained on 32T+ tokens with FP4/FP8 mixed precision; introduces Compressed Sparse Attention (CSA) using only 27% single-token FLOPs vs. V3.2 and 10% KV cache; scores 90.1 on GPQA Diamond and 80.6% on SWE-bench Verified. Open-weight.836
78
talkie-1930-13bIndependenthttps://talkie-lm.com/chat13Dense26020:10.2history onlyApr/2026🟢Ahttps://talkie-lm.com/introducing-talkieAlec Radford (GPT-1, GPT-2, GPT-3). talkie-1930-13b is a 13b language model trained on pre-1931 English-language text, instruction-tuned using a novel instruction-following dataset built from pre-1931 reference works including etiquette manuals, letter-writing manuals, encyclopedias, and poetry collections. It has also undergone reinforcement learning using online DPO to improve instruction-following capabilities.835
79
Hy3 previewTencenthttps://huggingface.co/tencent/Hy3-preview29521MoE40,000136:111.587.4265.7630
synthetic, web-scale
Apr/2026🟢Chttps://github.com/Tencent-Hunyuan/Hy3-previewReasoning"Hy3 preview is the first model trained on our rebuilt infrastructure, and the strongest we've shipped so far. It improves significantly on complex reasoning, instruction following, context learning, coding, and agent tasks."834
80
Ling-2.6-1TInclusion AIhttps://huggingface.co/inclusionAI/Ling-2.6-1T100050MoE20,75021:115.2
synthetic, web-scale
Apr/2026🟢Chttps://arxiv.org/abs/2606.15079SOTA"excellence across reasoning, coding, and tool-calling, achieving open-source SOTA status on multiple execution-heavy benchmarks:"833
81
GPT-5.5OpenAIhttps://chatgpt.com/3000150MoE200,000*67:181.693.657.2
synthetic, web-scale
Apr/2026🟢Dhttps://deploymentsafety.openai.com/gpt-5-5/gpt-5-5.pdfReasoning, SOTAAnnounce: https://openai.com/index/introducing-gpt-5-5/ HLE result is for GPT-5.5 Pro/x-high.832
82
Marul V7Independenthttps://marulai.com.tr/0.258Dense1,0003,876:10.05web-scaleApr/2026🟢Chttps://www.reddit.com/r/LocalLLaMA/comments/1sshwtu/s%C4%B1f%C4%B1rdan_e%C4%9Fitilmi%C5%9F_258m_parametre_t%C3%BCrk%C3%A7e_llm/Turkish.831
83
Qwen3.6-27BAlibabahttps://huggingface.co/Qwen/Qwen3.6-27B27Dense40,0001,482:13.586.187.824.3
synthetic, web-scale
Apr/2026🟢Chttps://huggingface.co/Qwen/Qwen3.6-27BReasoning"the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience."830
84
MiMo-V2.5-ProXiaomihttps://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro102042MoE27,00027:117.589.468.566.748
synthetic, web-scale
Apr/2026🟢Ahttps://mimo.xiaomi.com/mimo-v2-5-proReasoningUses a 7:1 Hybrid Attention mechanism and supports a 1M-token context window. "significant improvements over its predecessor, MiMo-V2-Pro, in general agentic capabilities, complex software engineering, and long-horizon tasks."829
85
Ling-2.6-FlashInclusion AIhttps://huggingface.co/inclusionAI/Ling-2.6-flash1047.4MoE20,750200:14.9
synthetic, web-scale
Apr/2026🟢Chttps://arxiv.org/abs/2606.15079ReasoningMoE with 104B total / 7.4B active params using a 1:7 MLA + Lightning Linear hybrid attention architecture for up to 4x throughput vs. comparable models; supports 262K context; achieves 61.2% on SWE-bench Verified and 73.85% on MathArena AIME 2026. Open-weight.828
86
Granite-4.1-8BIBMhttps://huggingface.co/ibm-granite/granite-4.1-8b8Dense15,0001,875:12.373.8455.9941.96
synthetic, web-scale
Apr/2026🟢Ahttps://huggingface.co/ibm-granite/granite-4.1-8bReasoning"improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities."827
87
OpenMythosIndependenthttps://github.com/kyegomez/OpenMythos0.770.04MoE3039:118.4web-scaleApr/2026🟢Ahttps://github.com/kyegomez/OpenMythos770M trained, up to 1T available. "OpenMythos is an open-source, theoretical implementation of the Claude Mythos model. It implements a Recurrent-Depth Transformer (RDT) with three stages: Prelude (transformer blocks), a looped Recurrent Block (up to max_loop_iters), and a final Coda. Attention is switchable between MLA and GQA, and the feed-forward uses a sparse MoE with routed and shared experts ideal for exploring compute-adaptive, depth-variable reasoning."826
88
Kimi K2.6Moonshot AIhttps://huggingface.co/moonshotai/Kimi-K2.6100032MoE30,50031:118.490.554
synthetic, web-scale
Apr/2026🟢Ahttps://www.kimi.com/blog/kimi-k2-6Reasoning, SOTA"Kimi K2.6 is an open-source, native multimodal agentic model"825
89
Qwen3.6-Max-PreviewAlibabahttps://chat.qwen.ai/2000100MoE40,000*20:120.0
synthetic, web-scale
Apr/2026🟢Dhttps://qwen.ai/blog?id=qwen3.6-max-previewReasoning"Qwen3.6-Max-Preview is an early preview of our next proprietary model, delivering meaningful improvements over Qwen3.6-Plus in agentic coding, world knowledge, and instruction following. It achieves the top score on six major coding benchmarks — SWE-bench Pro, Terminal-Bench 2.0, SkillsBench, QwenClawBench, QwenWebBench, and SciCode — with substantial gains over its predecessor. It also demonstrates stronger knowledge (SuperGPQA, QwenChineseBench) and better instruction following (ToolcallFormatIFBench)."824
90
Qwen3.6-35B-A3BAlibabahttps://huggingface.co/Qwen/Qwen3.6-35B-A3B353MoE40,0001,143:13.985.28621.4
synthetic, web-scale
Apr/2026🟢Chttps://qwen.ai/blog?id=qwen3.6-35b-a3bReasoning"Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6."823
91
Grok 4.3xAIhttps://grok.com/50025MoE80,000160:121.1
synthetic, web-scale
Apr/2026🟢Chttps://grok.com/release-notesReasoning"Grok 4.3 is a new pre-trained model matching the scale of Grok 4.20 with an improved architecture and a December 2025 knowledge cutoff." "0.5T total. Current Grok [4.2] is half the size of Sonnet and 1/10th the size of Opus." https://x.com/elonmusk/status/2042123561666855235 "The public facing v4.2 is based on foundation model v8, trained on Hoppers, with significant shortfalls in training data quality, comprehensiveness and proportionality. It is also only 0.5T in size." https://x.com/elonmusk/status/2055298325994164377822
92
Claude Opus 4.7Anthropichttps://claude.ai/ 5000150MoE80,000*16:166.794.254.7
synthetic, web-scale
Apr/2026🟢Dhttps://cdn.sanity.io/files/4zrzovbb/website/037f06850df7fbe871e206dad004c3db5fd50340.pdfReasoning, SOTAAnnounce: https://www.anthropic.com/news/claude-opus-4-7821
93
GPT-RosalindOpenAI3000150MoE114,000*38:161.6
synthetic, web-scale
Apr/2026🔴Fhttps://openai.com/index/introducing-gpt-rosalind/Reasoning"our frontier reasoning model built to support research across biology, drug discovery, and translational medicine. The life sciences model series is optimized for scientific workflows, combining improved tool use with deeper understanding across chemistry, protein engineering, and genomics."820
94
GPT-5.4-CyberOpenAI3000150MoE114,000*38:161.6
synthetic, web-scale
Apr/2026🔴Fhttps://openai.com/index/scaling-trusted-access-for-cyber-defense/Reasoning"In preparation for increasingly more capable models from OpenAI over the next few months, we are fine-tuning our models specifically to enable defensive cybersecurity use cases, starting today with a variant of GPT‑5.4 trained to be cyber-permissive: GPT‑5.4‑Cyber. "819
95
Marco-MiniAlibabahttps://huggingface.co/AIDC-AI/Marco-Mini-Instruct17.30.86MoE36,0002,081:12.683.470.750.3
synthetic, web-scale
Apr/2026🟢Chttps://huggingface.co/AIDC-AI/Marco-Mini-InstructReasoning"a highly sparse MoE multilingual model... 256 experts, 8 active per token. Drop-Upcycling from Qwen3-0.6B-Base. 2-stage post-training: SFT + Online Policy Distillation" Dataset: ""Web-mined NLLB translation pairs... Wikidata-sourced content with Gemini3-Flash text synthesis for cultural concepts."818
96
EXAONE 4.5LGhttps://huggingface.co/LGAI-EXAONE/EXAONE-4.5-33B33Dense14,000425:12.383.380.513.6web-scaleApr/2026🟢Chttps://arxiv.org/abs/2604.08644Reasoning“EXAONE”=“EXpert AI for EveryONE”. Training tokens/ratio: EXAONE-3 7.8B=8T tokens (Aug/2024) -> EXAONE-3.5 7.8B=9T -> EXAONE-3.5 32B=6.5T tokens -> EXAONE 4.0 32B=14T tokens. "We introduce EXAONE 4.5, the first open-weight vision language model developed by LG AI Research. Integrating a dedicated visual encoder... EXAONE 4.5 features 33 billion parameters in total."817
97
Muse SparkMeta AIhttps://meta.ai/50050MoE40,000*80:114.989.558.4
synthetic, web-scale
Apr/2026🟢Dhttps://ai.meta.com/blog/introducing-muse-spark-msl/Reasoning, SOTA"This initial model is small and fast by design, yet capable enough to reason through complex questions in science, math, and health." Announce: https://about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs/ "we can reach the same capabilities with over an order of magnitude less compute than our previous model, Llama 4 Maverick. This improvement also makes Muse Spark significantly more efficient than the leading base models available for comparison."816
98
Horus 1.0 4BTokenAIhttps://huggingface.co/tokenaii/horus4Dense3,000750:10.4856020
synthetic, web-scale
Apr/2026🟢Chttps://tokenai.cloud/horusReasoningFirst open-source AI model from Egypt.815
99
Ternary Bonsai 8BPrismMLhttps://huggingface.co/collections/prism-ml/ternary-bonsai8.19Dense36,0004,396:11.8
synthetic, web-scale
Apr/2026🟢Ahttps://github.com/PrismML-Eng/Bonsai-demo/blob/main/ternary-bonsai-8b-whitepaper.pdf1.58-bit ternary {-1,0,+1} weights across embeddings, attention, MLP, and LM head; quantization-aware retrain of Qwen3-8B. 9.4x smaller than FP16; 27 tok/s on iPhone 17 Pro Max.814
100
Claude Mythos Preview
Anthropic10000150MoE250,001*26:1166.794.564.7
synthetic, web-scale
Apr/2026🔴Fhttps://www-cdn.anthropic.com/78073f739564e986ff3e28522761a7a0b4484f84.pdfReasoning, SOTAClaude Mythos is suspected to be a Recurrent-Depth Transformer (RDT) — also called a Looped Transformer (LT). https://github.com/kyegomez/OpenMythos "Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available."813