Model Library/Tencent Hunyuan 3 (Hy3)

Tencent Hunyuan 3 (Hy3)

LLM
Chat

Tencent's 295B Mixture-of-Experts LLM (21B active) with strong agent and reasoning capabilities and 256K context.

On-Demand Dedicated 8xH200

Details

Modalities

text

Recommended Hardware

8xH200

Estimated Price

Loading...

Provider

tencent

Family

Hy3

Parameters

299B

Context

262144 tokens

License

apache-2.0

Tencent Hunyuan 3 (Hy3): Advanced Mixture-of-Experts LLM

Hy3 is a large Mixture-of-Experts (MoE) language model developed by the Tencent Hy Team. It activates only a small fraction of its experts per token (21B active parameters, with 192 experts and top-8 routing), pairing the quality of a very large model with the inference efficiency of a far smaller dense one. Following the Hy3 Preview release, the team scaled up post-training with higher-quality data gathered from feedback across 50+ products, producing a model that rivals flagship open-source models several times its active size and delivers strong gains on real-world productivity tasks. Weights are published on Hugging Face, with an FP8 checkpoint available at Hy3-FP8.

Key Features

  • Mixture-of-Experts efficiency - Sparse top-8 routing over 192 experts keeps active compute low while retaining large-model quality.
  • Long context - Native 256K-token context window for long documents, codebases, and extended multi-turn dialogue.
  • Strong agentic and reasoning capability - Post-trained with scaled reinforcement learning for reasoning, tool use, and long-horizon tasks.
  • Production-grade tool calling - Reliable tool-call and output-format handling that generalizes across agent scaffoldings.
  • Reasoning parser support - Ships with dedicated reasoning and tool-call parsers for vLLM and SGLang.

Architecture

Hy3 uses a Mixture-of-Experts transformer with 80 layers and grouped-query attention (64 attention heads, 8 key-value heads). A dedicated multi-token-prediction layer supports speculative decoding for lower latency. Only the top-8 of its 192 experts are activated per token, so the model reasons with a large knowledge capacity while keeping per-token computation modest.

Agentic and Reasoning Strengths

Building on Hy3 Preview, the team improved the quality and diversity of post-training data while scaling up reinforcement learning. Hy3 shows solid gains across reasoning, agentic, and long-context evaluations, remaining competitive with much larger flagship models. In a blind evaluation run with 270 domain experts using tasks drawn from their own work, Hy3 scored 2.67 out of 4, ahead of GLM-5.1 at 2.51, with its largest advantages in frontend development, data and storage, and CI/CD tasks. In productivity scenarios such as coding, office work, financial modeling, frontend design, and game development, Hy3 performs as a reliable, cost-effective option.

Reliability Improvements

Hy3 targets the operational failure modes that matter in production. Tool-call and output-format reliability were raised to production-grade standards, with error recovery and efficiency improved and cross-scaffolding variance on SWE-Bench Verified held within about 4 percent. An anti-hallucination training regime lowered the internal hallucination rate from 12.5 to 5.4 percent and commonsense error rates from 25.4 to 12.7 percent. Joint SFT and RL optimization improved coreference resolution, ellipsis recovery, and multi-turn constraint tracking, cutting the internal multi-turn issue rate from 17.4 to 7.9 percent and improving long-dialogue performance while keeping outputs concise.

Use Cases

  • Autonomous agents and tool-using assistants
  • Long-context document and codebase analysis
  • Coding, frontend development, and software engineering workflows
  • Multi-turn conversational assistants and customer support
  • Reasoning-heavy research, analysis, and productivity tasks

Deploy Tencent Hunyuan 3 (Hy3) on Vast.ai with vLLM or SGLang for scalable, OpenAI-compatible inference.

Quick Start Guide

Choose a model and click 'Deploy' above to find available GPUs recommended for this model.

Rent your dedicated instance preconfigured with the model you've selected.

Start sending requests to your model instance and getting responses right now.