สแนปชอตลงวันที่ 2 กันยายน 2026 ของ Qwen3.8 Max โมเดลเรซันนิงมัลติโมดัลเรือธงของ Alibaba: รับอินพุตข้อความ + รูปภาพ + วิดีโอ ส่งออกเป็นข้อความ รองรับบริบท 1M โทเคน มีความสามารถและราคาเดียวกันกับโมเดลฐาน สามารถปักหมุดเพื่อให้ได้ผลลัพธ์ที่ทำซ้ำได้
Qwen3.8 Max (0902) is a model entry on OrcaRouter from provider qwen, published with the model ID qwen/qwen3.8-max-0902. The notation 0902 is shown in the listing, but the supplied facts do not…
The model supports a context window of 1,000,000 tokens. That is the total amount of input the model can attend to in one request, including the text, image, and video content in the prompt. No further limit is given in the facts, so practical limits depend on how OrcaRouter and the provider enforce tokenization and request timeouts. For long text, 1,000,000 tokens allows many substantial documents to be included in a single prompt, reducing the need to chunk or summarize before calling the model. For video, the facts do not state how many tokens are consumed per second, per frame, or per video file, so you should estimate from metadata or test calls. It is also important to remember that the context window size does not tell you how accurate the model is on a 1,000,000-token prompt; no benchmark data is supplied. If your application needs exact long-context quality, evaluate on a sample of your own documents before committing to production.
With text, image, and video accepted as inputs, the model can be asked to reason over source material that combines those types in the same request. Examples include reading instructions in text, examining an image, and referring to a video when producing an answer. The provider fact sheet does not list task-specific capabilities such as OCR, object detection, or temporal video reasoning, so those behaviors should not be assumed. What can be expected from the model depends mainly on its trained abilities, which are not documented in the provided facts. The safe design is to use it for tasks that call for a text output from a mixed-modality prompt and that can tolerate the cost of storing images and video as input tokens. If a workload is text-only, another model without image and video input might be simpler. If a task needs exact timestamp-level video operations, the model card gives no evidence whether such behavior is reliable.
Choosing a cheaper model makes sense when the task does not require the features that define this listing. A short text question does not need a 1,000,000-token context or a 131,072-token output ceiling. A text-only workflow does not need image or video input. OrcaRouter bills this model at $2.00 per 1,000,000 input tokens and $6.00 per 1,000,000 output tokens. If a cheaper alternative has lower per-token prices and adequate context and modality support, it will reduce cost. There are no quality or speed benchmarks in the supplied facts, so choosing this model over other models cannot be justified by measured accuracy; it should be justified by explicit product requirements: large context, mixed text/image/video input, or long output in one generation. Because the input price applies to every token in the context, very large context calls can cost more than smaller calls even when the user question is short, so avoid padding prompts.
The supplied model facts for Qwen3.8 Max (0902) do not include benchmark scores, evaluation sets, or accuracy numbers. For that reason, any statement about model quality would be speculation, and OrcaRouter does not publish inferred scores. If you need a number to compare candidates, run your own tests on representative inputs, including the image and video cases you expect to send. A benchmark such as a public knowledge test would not tell you how well the model handles a specific 1,000,000-token video prompt. Keep in mind that a model name like Max is a label, not evidence of quality. The measurable things in this listing are the context window, maximum output length, input modalities, and price. Those attributes affect whether a workload fits, but they do not describe accuracy, reasoning depth, instruction following, or safety behavior. Treat capability in those areas as unverified unless you carry out an evaluation.
The model facts do not give tokens-per-second rates, time-to-first-token figures, request timeout guidance, or any other performance measurement. OrcaRouter does not claim a latency figure for this provider model. Speed will depend on the provider's serving infrastructure, the size of the input, and the amount of output requested. A 1,000,000-token input and a 131,072-token output are likely to take longer than a short request would, but the listing gives no exact relationship. Video input can also make latency worse because it has to be processed into tokens before the model can attend to it; the facts do not quantify that step. Reviewers should not assume that low per-token pricing implies fast decoding. The only speed-related decision supported by the facts is to test representative request sizes under your expected load. Do not infer real-time suitability from the context window size or the model name.
Stated strengths are structural. The model has a 1,000,000-token context window, a high 131,072-token maximum output, and accepts text, image, and video inputs. These are useful for applications that need one model call to observe a large or mixed body of material before writing a long response. Limitations are also easier to state from the facts than strengths. There are no quality benchmarks published, so accuracy, reasoning, and safety are undocumented. There is no separate listing for the training data, release notes, or update methodology behind the 0902 label. Image and video input do not guarantee exact visual or temporal comprehension. Context and output limits are maximums, not recommended usage; very large outputs may be expensive at $6.00 per 1,000,000 output tokens. A request that fills the 1,000,000-token context would consume $2.00 of input tokens before any output is generated, which is a meaningful operational cost.
The listed unit prices are $2.00 for every 1,000,000 input tokens and $6.00 for every 1,000,000 output tokens. OrcaRouter bills the model at the provider rate with zero markup, so the stated price is the provider rate passed through. Input tokens include the text, image, and video tokens sent in the prompt. Output tokens are generated response tokens. At these rates, 1,000,000 input tokens cost $2.00; 1,000,000 output tokens cost $6.00. A smaller request costs proportionally less: 100,000 input tokens would be $0.20 and 100,000 output tokens would be $0.60, provided those amounts are measured by the provider. The model card does not include separate pricing for cached input, batched requests, or interactive tiers, so no such prices should be assumed. Exact billed token counts should be checked from the OrcaRouter usage response or logs for each call.
When OrcaRouter says a model is billed at the provider rate with zero markup, it means OrcaRouter is not adding its own per-token fee on top of the provider rate for this listing. The final amount for token usage should therefore be based on the stated $2.00 input and $6.00 output prices per 1,000,000 tokens. That does not necessarily mean the final invoice has no other charges; taxes or account-level fees could apply depending on your arrangement, but no such fees are documented in these facts. It also does not mean the model is free to serve, because the per-token rate still applies. When comparing providers, the price shown by OrcaRouter for this model should represent the same units as this listing. If a different gateway quotes another price, the difference is billing structure, not a change in the model's context window.
Cost management should begin with prompt length. Since input costs $2.00 per 1,000,000 tokens, each additional token increases the input charge, and every image or video sent in a request is converted into tokens using rules not provided in this listing. Trimming irrelevant source material can reduce cost. Output is three times the input unit price, so limiting the requested output length also controls cost. A model with a 131,072-token output limit can generate responses that cost money; at $6.00 per 1,000,000 tokens, 131,072 output tokens would cost about $0.79 in output tokens alone. Generating that many tokens is not automatic, because the prompt controls response size. There is no cache price published for this model, so do not assume repeated context gets discounted. The absence of cache or batch pricing in the facts is not proof that such options do not exist; it simply means they are not part of this listing.
Qwen3.8 Max (0902) is exposed through OrcaRouter's OpenAI-compatible API. Set the client base URL to https://api.orcarouter.ai/v1 and use the exact model ID qwen/qwen3.8-max-0902. With an OpenAI-compatible SDK, the request structure is basically the same as any chat-completions call: pass an API key for OrcaRouter, the base URL above, and the model ID in the body. Do not use a generic name such as Qwen3.8 Max or qwen3.8; the listing requires qwen/qwen3.8-max-0902. The same model ID should be used in direct HTTP requests to the base URL, where the path and payload follow the OpenAI-compatible contract exposed by OrcaRouter. Because it is OpenAI-compatible, existing code written for OpenAI chat completions can usually be migrated by changing the base_url and model parameters, though authentication details must match OrcaRouter's requirements.
For an OpenAI-compatible chat completion, parameters such as model, messages, and output token limit are handled like other compatible models. The facts give a maximum output of 131,072 tokens, so if you set an output limit, it should be no higher than 131,072; the default output limit is not stated in the facts. Temperature, top-p, stop, and other sampling parameters are not documented in the supplied model facts, so follow OrcaRouter's API documentation and standard OpenAI parameter semantics. For image and video input, use the multimodal content format expected by OrcaRouter's API; the facts do not provide a sample. If the API returns an error about unsupported parameter names, check whether your SDK is sending legacy parameters. The context window is 1,000,000 tokens, so the prompt must keep all input tokens, including image and video tokens, below that limit. Responses count separately against the 131,072-token maximum.
Because OrcaRouter offers an OpenAI-compatible API at https://api.orcarouter.ai/v1, migration is mostly a configuration change. Replace the base URL with https://api.orcarouter.ai/v1, replace the API key field with an OrcaRouter key, and set model to qwen/qwen3.8-max-0902. If your code uses an OpenAI-compatible client library, update the client's base_url and api_key and leave most message structures intact, assuming the multimodal format matches this model. The facts do not say whether this model supports every OpenAI-specific parameter, so before migrating, review the parameters your application sends. Also, because the context limit is 1,000,000 and the maximum output is 131,072, adjust request truncation logic to those limits. Any existing code that hardcoded a different maximum context length may need to be updated. Finally, confirm that your application's data-handling expectations align with OrcaRouter's terms, because the model facts themselves do not discuss data retention.
The main basis for comparison is the input contract. Qwen3.8 Max (0902) accepts text, image, and video, so a text-only alternative would eliminate those modalities. If your actual workload contains only text, a text-only model with a large enough context is likely a more direct fit and may have different pricing. The model facts do not include accuracy or speed comparisons against any other model, so you cannot choose based on public quality scores. You can compare structural attributes: a text-only alternative may have a smaller context, a shorter output limit, or a lower input price. OrcaRouter bills this model at $2.00 input and $6.00 output per 1,000,000 tokens. The fact that it has image and video input is useful only if those inputs are used. If those inputs are not used, you may be paying for multimodal capacity that is irrelevant to the task.
Choose Qwen3.8 Max (0902) when the task profile matches the listed features. First, you need a 1,000,000-token context because source material cannot be shortened to fit a smaller window. Second, you need image and video input in the same prompt, or you expect those modalities in future requests. Third, you want an output maximum of up to 131,072 tokens. If your workload does not have any of those requirements, many of the benefits of this listing are unused. Do not choose it because the name Max appears strong; the listing contains no benchmark evidence. Since OrcaRouter exposes models through the same OpenAI-compatible API, swapping model IDs is easy, so you can test this model against another candidate on your own task. Just remember that input and output prices differ, and no cache pricing is specified for this model.
When comparing cost, first normalize to the same token basis. This model uses $2.00 per 1,000,000 input tokens and $6.00 per 1,000,000 output tokens, passed through from the provider at zero markup. A competing model at a lower per-token input price may actually be cheaper for a full context request, but context window and maximum output must be considered too. For example, a 1,000,000-token request can use this model's full context in one call; a model with a 200,000-token context would require multiple calls, and the additional output or summarization passes can change total price. The supplied facts do not provide objective accuracy or latency values, so any cost-per-correct-answer comparison must come from your own tests. Also check whether a competing model imposes separate charges for image or video tokens, because those are not described here. When in doubt, run a typical workload on OrcaRouter and compare measured token usage and response quality rather than relying only on list prices.
เข้ากันได้กับ OpenAI — ใช้ SDK เดิมของคุณได้เลย
https://api.orcarouter.ai/v1import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.orcarouter.ai/v1",
api_key=os.environ["ORCAROUTER_API_KEY"],
)
response = client.chat.completions.create(
model="qwen/qwen3.8-max-0902",
messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)enable_thinkinginclude_reasoningmax_tokenspresence_penaltyreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p| อินพุต / 1M โทเค็น | $2.00 |
| เอาต์พุต / 1M โทเค็น | $6.00 |
| อ่านแคช / 1M | $0.250 |
| เขียนแคช / 1M | $2.50 |
| สกุลเงิน | USD |
ประมาณการจากราคาตั้ง
เป็นเพียงการประเมิน — จำนวน token จริงขึ้นอยู่กับ tokenizer ของผู้ให้บริการ
สิ่งที่นักพัฒนาพูดถึงในสัปดาห์นี้
GET /api/public/models/qwen/qwen3.8-max-0902เปิด @misc{orcarouter_qwen3_8_max_0902,
title = {Qwen3.8 Max (0902) API},
author = {Qwen},
year = {2026},
howpublished = {OrcaRouter},
url = {https://www.orcarouter.ai/models/qwen/qwen3.8-max-0902}
}Qwen. (2026). Qwen3.8 Max (0902) API. OrcaRouter. https://www.orcarouter.ai/models/qwen/qwen3.8-max-0902