A hero title card for Apodex 1.1 titled 'Apodex 1.1' with the subtitle '44 on the Intelligence Index - agentic strength, heavy verbosity, and a knowledge catch', pill badges reading 'AA Index 44', '#11 of 177', '256K context', and a footer line 'Released Aug 24, 2026 - first Apodex model on Artificial Analysis', with the OrcaRouter logo in the bottom-right corner.
Guides & Insights

Apodex 1.1, explained: a 44 on the Intelligence Index, real agentic strength — and a knowledge-reliability catch

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Apodex 1.1 is a week old, proprietary, and until this week it was mostly a lab curiosity. Then Artificial Analysis published its first independent evaluation of the model — Apodex's first on the platform — and the scoreboard did something unusual: Apodex 1.1 landed a 44 on the Artificial Analysis Intelligence Index, ranked #11 of 177, while posting agentic-work scores that beat every model in its intelligence tier, including DeepSeek V4 Pro, Qwen3.7 Max and Kimi K2.6. The same report surfaced a second, uglier number: a knowledge-reliability record weak enough to change the decision about running real work on it. Its open-weights sibling, Apodex 1.1 Mini, is the model most people can actually run. Here is what the evaluation says, what it doesn't, and who should care.

Apodex, the AI company founded by Chen Tianqiao, shipped the family on August 24, 2026, alongside a 70-author arXiv paper, "Apodex 1.1: Scaling Agentic Intelligence for Complex Work" (2608.23283). The flagship is proprietary — roughly 397B parameters, per the vendor's report — built around the thesis that the real bottleneck for agents is not raw intelligence but the environment they execute in. Apodex 1.1 was therefore trained to work directly with files, search and code over long horizons, with a shared execution harness ("AgentOS") that coordinates up to 150 parallel sub-agents and keeps a provenance record of every step. The open part is the sibling: Apodex 1.1 Mini, a 35B-class MoE fine-tuned from Alibaba's Qwen3.5-35B-A3B, released under Apache-2.0 with a companion agent harness called FrontierAgent. The full model reaches you only through Apodex's own API and online workbench; the Mini is self-hosted or nothing.

What a 44 on the Intelligence Index means

The Artificial Analysis Intelligence Index is a weighted aggregate of the platform's benchmark suite — including agentic evals like GDPval-AA v2 and Terminal-Bench, plus reasoning, math and knowledge components. A score of 44 puts Apodex 1.1 at #11 of 177 models, well above the median of 18. It also drops the model into a crowded, interesting tier: Kimi K2.6 (45), MiniMax-M3 (45) and Inkling (42) all sit within three points, and Apodex 1.1 is the only one of that group whose name was unknown to most AI users a week ago. Artificial Analysis notes that the page covers the reasoning version of the model, and that a non-reasoning variant may also exist.

A single-column scoreboard for Apodex 1.1 titled 'Apodex 1.1 - the scoreboard' listing AA Intelligence Index 44 (#11 of 177), GDPval-AA v2 Elo 1348, TerminalBench v2.1 70%, Output tokens per task ~17,000, Full Index evaluation cost 18.41, and AA-Omniscience -21.9 (78% hallucination), with a footer 'All figures per Artificial Analysis, August 2026.' and the OrcaRouter logo in the bottom-right corner.

The headline index score is not where this model gets interesting. It is interesting because of where its strengths and weaknesses sit relative to that score — and that is what the tier comparison makes concrete.

Where it punches above its tier: agentic and knowledge-work evals

On the two Index components that measure real-world agentic work, Apodex 1.1 over-performs its 44-point neighbors. On GDPval-AA v2 — a benchmark that scores an agent's ability to complete realistic office-style tasks end to end — Apodex 1.1 posts an Elo of 1348, ahead of DeepSeek V4 Pro (1333), Qwen3.7 Max (1308) and Kimi K2.6 (1202). On TerminalBench v2.1, which tests an agent working at a terminal, it scores 70%, ahead of Kimi K2.6 (66%) and just behind Qwen3.7 Max (75%). Those are independent, Artificial Analysis-run numbers, and they are the strongest evidence in the report that the environment-first training works.

The vendor's own benchmark claims go further, and should be read as vendor-reported until someone reproduces them: APEX-Agents 38.5, GDPVal 78.8, SWE-bench Verified 77.7%, Terminal-Bench 2.1 70.8%, Humanity's Last Exam 56.1% with tools, DeepSearchQA 92.4%, FrontierFinance 54.3, and FrontierScience-Research 63.3. What makes the agentic profile credible rather than hype is the architecture behind it: environment scaling (expanding the diversity of executable file, search and code environments used in training) and agentic coordination scaling (training the model to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results and replan). "Agent Team" mode spawns up to 150 sub-agents on retrieval and synthesis tasks, and a built-in verification step ("Statement Review") checks conclusions before they are delivered. In other words: this is a model designed to be wrong, check itself, and fix things — not to recite facts from memory.

The hidden cost of all that agency: verbosity

That design shows up on Artificial Analysis's cost-efficiency table as the model's clearest weakness after knowledge: it is very verbose. Running the full Intelligence Index consumed 180M output tokens — versus a 67M median — and roughly 17,000 output tokens per benchmark task, against about 9,400 for Qwen3.7 Max and 8,100 for MiniMax-M3. In total, the full evaluation cost $618.41 in API spend, per Artificial Analysis.

A screenshot of the Artificial Analysis model page for Apodex 1.1 (captured August 31, 2026), showing 'Apodex 1.1 scores 44 on the Artificial Analysis Intelligence Index' with a median of 18, the note that it generated 180M tokens versus a 67M median, pricing of $0.30 per 1M input and $3.00 per 1M output tokens, and the $618.41 total cost to evaluate the model on the Intelligence Index.

The pricing that produces that bill: $0.30 per million input tokens and $3.00 per million output — a $0.03 cache-hit price on cached input, a 90% discount — which blends to about $0.38 per million on a typical 7:2:1 cache/input/output mix. Both per-token prices sit above the platform medians ($0.25 input, $0.90 output). Yet Artificial Analysis's cost-per-task estimate still puts Apodex 1.1 at roughly $0.05 per benchmark task, cheaper than Qwen3.7 Max (~$0.07) and Kimi K2.6 (~$0.06). The reconciliation matters for budgeting: its tokens are cheap enough that even heavy verbosity stays competitive per task — the cost risk is volume, not rate. If you put a long-running agent on it, the bill is driven by how much it talks, and it talks a lot.

One pricing note before you budget: $0.30/$3.00 is the vendor's own list. A routing platform that passes provider list prices through at 0% markup charges exactly that figure, and any future Apodex price cut shows up the same day it is announced.

The catch: a weak knowledge-reliability record

Here is the number the launch coverage mostly misses. On AA-Omniscience, Artificial Analysis's knowledge-reliability probe, Apodex 1.1 scores -21.9. It attempts 87% of questions rather than declining, but gets only 32% right, and its hallucination rate is 78.4%. For a model positioned as a heavy-duty solver, that is a serious split personality: strong on agentic execution, weak on grounded knowledge.

The two facts are connected, not contradictory. Agentic benchmarks reward acting in an environment — issuing tool calls, iterating, recovering — so a model can score well there while having poor recall, because the environment does the remembering. The design even assumes it: Statement Review exists to check outputs, and the workbench is built around verification, not around the model being right from memory. The practical reading is blunt: do not point Apodex 1.1 at tasks that depend on it retrieving facts from its own weights. Point it at long-horizon tasks where its work can be checked against files, code and sources — and verify the delivery.

The open sibling: Apodex 1.1 Mini and FrontierAgent

If the flagship's closed door is a problem, the family's open half is the counterweight. Apodex 1.1 Mini is a ~35B-total / ~3B-active MoE fine-tuned from Qwen3.5-35B-A3B (a base Alibaba open-sourced in February 2026), released under Apache-2.0 with a 262K context window and FP8, FP4 and GPTQ-Int4 variants on Hugging Face. Its vendor-reported headline is 50.2 on FrontierFinance (first in Apodex's own comparison) and 27.7 on APEX-Agents with the Agent Team harness enabled. And FrontierAgent — the open-source runtime that comes with it — is genuinely the interesting artifact: a terminal-based harness with ReAct and Agent Team modes, sandboxed file work, approval gates, checkpoints and traces, all talking to any OpenAI-compatible endpoint.

Two things worth knowing before you self-host. First, the Mini is a fine-tune, not a frontier model — read its benchmark numbers the same way you read the flagship's vendor table: unreproduced, harness-included, vendor-chosen comparisons. Second, its behavior inherits the base: if you want to test whether the underlying Qwen3.5-35B-A3B already does your task, you do not need a GPU and a serving stack to find out — the base is one API call away.

Who should use Apodex 1.1 today

The flagship is early, closed, and only on the vendor's own API and online workbench for now; Apodex says broader API availability is coming through major platforms, but the weights are not open. That constrains who can act on this week's scoreboard: teams already on the Apodex workbench, or willing to sign up for a first-party API key. The Mini is the more accessible path, but it means self-hosting.

A screenshot of the Apodex online workbench homepage (captured August 31, 2026), showing the deep-research interface with the prompt field 'Ask the question that matters', feature points for a step-level reasoning trace, citations in every report, branching off any report, and full-text search across threads, and a sign-in prompt for saving inquiries.

Where the model genuinely earns its place: long-horizon agentic pipelines where outputs are verified downstream — research workbenches, multi-step file-and-tool tasks, terminal work — and where the ~$0.05-per-task cost profile survives the verbosity. Where it does not belong yet: anything that depends on the model's own knowledge being trustworthy, or a production path that cannot tolerate an unproven seven-day-old model misbehaving. For exactly that situation, a routing layer is the right wrapper. One API for 200+ models means the base Qwen3.5-35B-A3B is already callable at list price, a self-hosted Mini can sit behind the same router with automatic failover, and a flaky newcomer cannot take a production path down — when it underperforms, the request rolls to a healthy provider instead.

What to watch next

This evaluation is a first reading, not a verdict. Three things will decide whether Apodex 1.1 matters in a month: independent reproduction of the vendor's agentic claims (the Artificial Analysis-run GDPval and TerminalBench numbers are the only ones not produced by Apodex itself); whether the knowledge-reliability gap narrows in a non-reasoning variant or a follow-up release; and whether the model shows up on more APIs and more independent scoreboards, at which point the 44 and the -21.9 become someone else's problem to beat. For a seven-day-old model with this kind of split profile, that is about as much as you can ask: a real agentic edge, an honest price, and one weakness you know about before you sign up.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily