Hy4 preview (Free)

tencent/hy4-preview-free
FREE
ToolsJSONReasoning
by Tencent

Hy4 preview is Tencent latest Mixture-of-Experts flagship: 770B total parameters with 49B activated per token, 78 layers, 256 routed experts plus 1 shared expert, and a 1M-token context window. It is a text-only model built for coding agents, complex tool-use workflows, and productivity work, with a native MTP layer for speculative decoding. Reasoning is configurable at three levels (high, low, none) and defaults to high, so it can run as a deep chain-of-thought model for math, coding and analysis, or answer directly when latency matters. It supports native tool calling, structured outputs, and the full sampling parameter set; Tencent recommends temperature 0.9 with top_p 1.0. Tencent positions it for software engineering, office and data analysis, game prototyping, and scientific research, and reports an internal blind evaluation (163 experts, 203 engineering tasks) in which Hy4 preview averaged 2.99 against GLM 5.3 at 2.92 and Kimi K3 at 2.94. As a preview release Tencent notes known rough edges, including spending longer than necessary on reasoning and over-verifying its own work.

ctx1M tokens
Max output64K
Inputtext
Outputtext
p50 TTFT5.25 s
PRICEFreerate-limited · model usage at $0
p50 TTFT5.25 s7d
p95 TTFT10.00 s7d
TRAFFIC52.9Mtokens / 7d

Code samples

import os

from openai import OpenAI

client = OpenAI(
    base_url="https://api.orcarouter.ai/v1",
    api_key=os.environ["ORCAROUTER_API_KEY"],
)

response = client.chat.completions.create(
    model="tencent/hy4-preview-free",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

Supported parameters

  • frequency_penalty
  • include_reasoning
  • max_completion_tokens
  • max_tokens
  • presence_penalty
  • reasoning
  • reasoning_effort
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_p

Pricing

$0
Pricing
Per request$0
BillingModel usage is never charged to your balance
Over the limitHTTP 429 when a limit is hit

Need it without the limits? tencent/hy4-preview·How free-tier limits work →

Performance

p50 TTFT
5.25 s
Output speed
77.0 tok/s
p95 TTFT
10.00 s
Error rate
1.4%

Public benchmarks

Source: Design Arena

Community buzz

What developers are saying this week

Hacker News0 mentions · 7d

FAQ

How much does Tencent: Hy4 preview (Free) cost on OrcaRouter?
Tencent: Hy4 preview (Free) is priced at $0.0000 per request via OrcaRouter (flat per-call fee, charged per generation rather than per token).
What is Tencent: Hy4 preview (Free)'s context window?
Tencent: Hy4 preview (Free) supports a context window of 1M tokens. Use long-context features (RAG, summarisation) up to that limit.
How do I call Tencent: Hy4 preview (Free) via the OpenAI SDK?
Set OpenAI base_url to https://api.orcarouter.ai/v1, supply your OrcaRouter API key, and pass model="tencent/hy4-preview-free" in the chat.completions.create call.
Does OrcaRouter rate-limit Tencent: Hy4 preview (Free)?
Per-model rate limits follow your OrcaRouter plan. Free tiers ship with conservative caps; paid tiers lift them. Check /pricing for current quotas.

Embed this badge

Tencent: Hy4 preview (Free)•pricing pending•5246ms p50•via OrcaRouter
HTML <a href="https://www.orcarouter.ai/models/tencent/hy4-preview-free" target="_blank"> <img src="https://www.orcarouter.ai/embed/tencent/hy4-preview-free.svg" alt="Tencent: Hy4 preview (Free) on OrcaRouter" /> </a>
Markdown [![Tencent: Hy4 preview (Free)](https://www.orcarouter.ai/embed/tencent/hy4-preview-free.svg)](https://www.orcarouter.ai/models/tencent/hy4-preview-free)

Model card as data

GET /api/public/models/tencent/hy4-preview-freeOpen
Machine-readable:/llms.txt/llms-full.txt