Trending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#8Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10Laguna-S-2.1poolside⬇3kTrending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Unlimited-OCRbaidu⬇2237kTrending Model:#4Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#7Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#8Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#9krea2-identity-editconradlocke⬇0kTrending Model:#10Laguna-S-2.1poolside⬇3k

AntAngelMed Deploys 100B Clinical MoE Model Locally in a Snap

A single minimalist clipboard with floating pen constructed from delicate translucent digital mesh hexagons and glowing strands.

AntAngelMed is a new open-source medical language model designed to assist with clinical reasoning and diagnosis. The model uses a mixture-of-experts architecture, activating only 6.1 billion of its 100 billion parameters for fast, efficient performance. It currently leads all open-source competitors on OpenAI’s HealthBench and several Chinese medical evaluation benchmarks.

The project was jointly developed by the Health Information Center of Zhejiang Province, Ant Healthcare, and Zhejiang Anzhen'er Medical AI Technology, with model weights hosted on Hugging Face by MedAIBase. They trained AntAngelMed through a three-stage pipeline that included continued pre-training on medical texts, supervised fine-tuning on high-quality instructions, and reinforcement learning to improve safety and empathy. The result is a small-activation model that can run on a range of hardware while matching much larger dense models.

Efficient MoE architecture and training

Key Features
  • Ranks first on OpenAI’s HealthBench benchmark.
  • Uses MoE with 6.1B active parameters.
  • Three-stage training: pre-training, fine-tuning, RL.
  • Supports 128K context length for documents.
  • Over 200 tokens per second on H20.
  • FP8 quantization and EAGLE3 speedups.

Privacy-sensitive medical professionals and clinics can deploy AntAngelMed locally, keeping patient data off cloud servers. Serious hobbyists with prosumer GPUs or Ascend hardware can experiment with its diagnostic capabilities for research. Small agencies building health chatbots can benefit from its openly licensed weights and extensive medical knowledge.

Developer notes on efficiency and hardware

The model activates only 6.1 billion parameters, enabling it to match the performance of dense models around 40 billion parameters. On H20 hardware it generates more than 200 tokens per second, roughly three times faster than a 36B dense model, and speed advantages grow with longer outputs. The team also applied FP8 quantization with EAGLE3 optimization, boosting throughput by up to 94% on some benchmarks under concurrent use.

"AntAngelMed surpasses all open-source models and a range of top proprietary models on OpenAI's HealthBench, and ranks first overall on the Chinese authority benchmark MedAIBench." — Source: Hugging Face