Model Library/Hy4 Preview

Hy4 Preview

LLM
MoE
Reasoning
Code
Chat

770B MoE flagship with 49B active parameters, Gated DSA attention, and 1M-token context

On-Demand Dedicated 8xB300

Details

Modalities

text

Recommended Hardware

8xB300

Estimated Price

Loading...

Provider

Tencent

Family

Hy4

Parameters

780B

Context

1048576 tokens

License

apache-2.0

Hy4 Preview: 770B Mixture-of-Experts Flagship with 1M-Token Context

Hy4 Preview is a new-generation Mixture-of-Experts flagship model from the Tencent Hy Team. It has 770B total parameters with 49B active per token and serves a 1M-token context window. Tencent describes it as an early release built by scaling model size, context length, and training data together, with a much larger post-training run than the previous generation, and reports the largest generation-over-generation gain the team has measured.

Key Features

  • Sparse Mixture-of-Experts - 256 routed experts and 1 shared expert, with the top 8 routed experts plus the shared expert active per token, so only 49B of the 770B parameters run on any given token
  • 1M-Token Context - Long-horizon development work, large-repository navigation, and extended agentic trajectories in a single context
  • Gated DeepSeek Sparse Attention - Sparse attention with an IndexCache that reuses the sparse index across layers, cutting long-context serving cost
  • Native Draft Layer - A multi-token-prediction layer ships in the checkpoint for speculative decoding
  • Adjustable Reasoning Effort - Reasoning effort defaults to high for deep chain-of-thought; a no-think setting returns direct answers
  • Domain-Trained on Expert Data - Training data was built with Tencent software engineers, game developers, finance analysts, and security experts, and the model is co-designed with Tencent's CodeBuddy and WorkBuddy products

Use Cases

  • Software engineering: understanding, planning, debugging, and verifying long-horizon development tasks, including front-end visual quality
  • Office and analysis work: turning scattered context into documents, spreadsheets, and presentations, and building equations and financial models
  • Game development: prompt-to-playable prototyping and multi-turn work against game engines
  • Scientific research across AI research, molecular dynamics, condensed matter physics, and pure mathematics
  • Agentic workflows with tool calling over very long contexts

Architecture and Design

The model has 78 layers. The first uses a dense feed-forward network and the other 77 are Mixture-of-Experts. Attention is Gated DeepSeek Sparse Attention across 64 heads, with query compression to 2048 dimensions and key-value compression to 512, and an indexer that selects the top 2048 positions. The residual pathway uses identity Hyper-Connections across four residual streams to widen information flow between layers. Hidden size is 6144 and the vocabulary is 120,832 tokens. Tencent credits DeepSeek and GLM as influences on the attention design.

Evaluation

Tencent ran a blind side-by-side study in which 163 internal experts rated model outputs on 203 engineering tasks. Hy4 Preview scored 2.99 against GLM 5.3 at 2.92, winning 46.8 percent of comparisons, and 2.99 against Kimi K3 at 2.94, winning 51.2 percent. Published benchmark results include 92.3 on GPQA Diamond, 85.4 on Terminal-Bench 2.1, 82.9 on SWE-bench Multilingual, 74.1 on Toolathlon Verified, 65.7 on SWE-bench Pro, 64.3 on Deep-SWE, and 62.9 on SkillsBench. Tencent labels this an early release with real headroom left, and calls out over-long reasoning on complex tasks and a tendency to over-verify its own work.

Deploy Hy4 Preview on Vast.ai for long-horizon software engineering, agentic tool use, and million-token context work on flexible GPU infrastructure.

Quick Start Guide

Choose a model and click 'Deploy' above to find available GPUs recommended for this model.

Rent your dedicated instance preconfigured with the model you've selected.

Start sending requests to your model instance and getting responses right now.