بطاقة عنوان رئيسية لمراجعة Qwen3.8-27B، بعنوان «مراجعة Qwen3.8-27B» مع العنوان الفرعي «نموذج 27B مفتوح الأوزان الذي يضاهي Opus في المنزل — مُختبر في مواجهة الضجيج»، تعرض شارة الأوزان المفتوحة، وشارة Apache 2.0، ومجموعة أيقونات GPU وخادم على خلفية بيضاء نظيفة مع لمسات تدرّج أزرق ناعم.
Guides & Insights

مراجعة Qwen3.8-27B: نموذج 27B مفتوح الأوزان الذي يُلقَّب بـ"أوبوس في المنزل" — اختبارٌ على محك الضجة

الكاتب

Alistair Wren

تاريخ النشر

أحدث النماذج · 20عرض جميع النماذج
المعايير: Artificial Analysis · يُحدَّث يوميًا
العودة إلى جميع المقالات

Qwen3.8 27B هو أفضل نموذج مفتوح الأوزان بحجم 27B يمكنك استضافته ذاتيًا في الوقت الحالي، ولا يوجد منافس قريب منه. صدرت الأوزان في 14 أغسطس 2026 بموجب ترخيص Apache 2.0: نموذج كثيف بحجم 27B (28B إذا حسبنا مشفر الرؤية) مع سياق أصلي يبلغ 262K، وإدخال صور وفيديو بشكل أصلي، ونتائج برمجة وكيلية معلنة من البائع تصل أو تتجاوز Cla​ude Opus 4.6 Max — SWE-bench Pro 61.7، وDeepSWE 1.1 42.2، وLiveCodeBench v6 90.3. السعر المعلن هو صفر: حمّل 55.6 غيغابايت وشغّله. لكن التحفظات حقيقية أيضًا — فالاختبارات المستقلة تجده أبطأ بنحو ثلاث مرات وأكثر استهلاكًا للرموز من سابقه، وكل المعايير في البطاقة لا تزال مبلّغة من Ali​baba، وسياق الـ 1M ميزة متاحة فقط في الاستضافة السحابية. الخلاصة: إذا كان لديك GPU بسعة 24 غيغابايت أو أكثر وتريد قدرة قريبة من الحدود القصوى دون فاتورة محسوبة، فاستضفه ذاتيًا — ولكن ادخل وأنت تعرف بالضبط أين تكون الأرقام هشّة.

الخلاصة أولاً: هل يجب أن تستخدم Qwen3.8 27B؟

إجابة مختصرة — نعم لأحمال العمل المحلية والخاصة؛ انتظر الأرقام من جهات خارجية قبل قرار الشراء. Qwen3.8 27B هو النموذج النادر الذي تكون فيه الأوزان المفتوحة هي المنتج والـ API هو وسيلة الراحة. نظرًا لأن الترخيص هو Apache 2.0، فإن سعر كل توكن هو صفر بشكل دائم بمجرد امتلاك العتاد، ولا يمكن إعادة ترخيص أي شيء تبنيه عليه أو فرض رسوم عليه بأثر رجعي. ما تتنازل عنه هو السرعة والتحقق والراحة.

إذا كنت تشغّل بالفعل Qwen3.6-27B وكنت راضيًا عنه، فإن الترقية حقيقية لكنها ليست مجانية من حيث الوقت الفعلي. إذا أتيت إلى هنا من إطلاق Qwen3.8-Max، فهذا هو العضو الأصغر القابل للاستضافة الذاتية من نفس الجيل — أداة مختلفة لمهمة مختلفة.

الحقائق السريعة، تم التحقق منها في 15 أغسطس 2026

صدر — أُطلقت الأوزان في 14 أغسطس 2026 على Qwen/Qwen3.8-27B على Hugging Face، وتمت مرآتها على ModelScope؛ وتشير التغطية المرتبطة بالمناطق الزمنية أيضًا إلى 13 أغسطس، وصدر الإصدار بعد حوالي أسبوع من الإعلان عن جيل Qwen3.8.

الترخيص — Apache 2.0: تنزيل، تعديل، إعادة توزيع، استخدام تجاري، مع منح براءة اختراع صريح. دائم.

المعلمات — كثيف بحجم 27B (28B عند احتساب مشفّر الرؤية)، 64 طبقة، حجم مخفي 5,120، مفردات 248,320.

البنية — انتباه هجين: 48 طبقة انتباه خطي من Gated DeltaNet مقابل 16 طبقة انتباه مقيدة كاملة (بنسبة 3:1). لهذا السبب يمكن لنموذج كثيف بحجم 27B أن يحمل 262K رمزًا من السياق الأصلي.

السياق — 262,144 رمزًا أصلًا؛ قابل للتوسيع إلى 1,000,000 عبر YaRN في النسخة المستضافة.

الإدخال — صورة وفيديو أصليان إلى جانب النص؛ يُرجع نصًا. مُرمّز الرؤية هو سبب بلوغ عدد المعلمات 28B.

التفكير — وضع التفكير مفعّل افتراضيًا ويمكن تعطيله عند الطلب؛ reasoning_effort (منخفض/متوسط/مرتفع) و preserve_thinking لتشغيلات الوكيل الطويلة.

الأوزان — 55.6 جيجابايت من ملفات safetensors بتنسيق BF16 في 18 شريحة؛ كما تتوفر إصدارات FP8 وGGUFs المجتمعية.

المواصفات أعلاه مأخوذة من بطاقة نموذج Hugging Face والإعدادات الخاصة بـQwen3.8 27B، التي اطلعت عليها اليوم. ادعاءات المعايير مقدمة من Ali​baba؛ وحتى اليوم لم يقم أي مختبر مستقل بتكرارها.

ما الذي يجيده Qwen3.8 27B حقًا

أقوى حالة هي هندسة البرمجيات الوكيلية. أعلنت Ali‌baba عن قفزة بمقدار 3 أضعاف في DeepSWE 1.1 مقارنة بالإصدار السابق 27B (42.2 مقابل 13.3) وبدرجة تفوق Cla​ude Opus 4.6 Max على SWE-bench Pro (61.7 مقابل 53.4). تلك هي الأرقام التي أكسبتها لقب "Opus at home" — نموذج 27B كثيف يقوم بأعمال برمجة قريبة من الحدود على أجهزة يمكن للشخص امتلاكها فعليًا.

Benchmark scoreboard card comparing Qwen3.8-27B with Qwen3.6-27B and Claude Opus 4.6 Max: SWE-bench Pro 61.7 vs 53.5 vs 53.4; DeepSWE 1.1 42.2 vs 13.3 vs n/a; LiveCodeBench v6 90.3 vs 83.9 vs 88.8; Terminal Bench 2.1 73.0 vs 63.4 vs 78.2; GPQA Diamond 89.2 vs 87.8 vs 91.3; OSWorld-Verified 84.3 vs 63.9 vs 72.7; QwenSWEBench 79.0 vs 49.3 vs 63.8; CoWorkBench 70.7 vs 61.0 vs 68.2. Footer: all figures Alibaba-reported, not yet independently reproduced.

ثلاثة أشياء إلى جانب الأرقام الخام تجعلها عملية شراء مبررة:

متعدد الوسائط أصليًا. صور وفيديو كمدخلات، ونصوص كمخرجات — المستند الذي يحتوي على مخططات أو شرح بالفيديو ليس نموذجًا منفصلًا. درجات الاستدلال البصري (85.6 مع تفعيل ميزة سلسلة التفكير، و94.6 في الرياضيات البصرية، وكلاهما مُعلَن من البائع) هي حيث يُثبت مُرمِّز الرؤية جدارته.

سياق أصلي 262K.مهام الوكلاء طويلة الأفق، وقواعد الأكواد الكبيرة، ونصوص المحادثات متعددة الساعات تتسع جميعها في نافذة واحدة. تصميم الانتباه الخطي الهجين يُبقي تكلفة KV لهذا السياق أقل من نموذج الانتباه الكامل النقي بنفس الحجم.

أوزان مجانية ودائمة. هذا هو جوهر الفئة: نموذج API مغلق يُستأجر؛ Qwen3.8 27B يُملَك. بدقة 4-بت يعمل على بطاقة رسومات بذاكرة 24 جيجابايت (من فئة RTX 3090/4090)، وقدَّمت AMD دعمًا من اليوم الأول (حتى 24.5 توكن/ثانية على Ryzen AI Max+ 395 و51.8 توكن/ثانية على Radeon AI PRO R9700، وفقًا لمدونة AMD الخاصة).

حيث يصبح الاستعراض قاسيًا: السرعة، والرموز، والتحقق

الاختبار المستقل حتى الآن هو إطار عمل لأحد المختبرين، وليس مختبرًا — لكنه أفضل إشارة لدينا. تشغيل Qwen3.8 27B ضد Qwen3.6-27B على عشر مهام واقعية (تم تقييمها بواسطة GPT-5.5 عبر إطار عمل llmcompare)، فاز إصدار 3.8 في تسع من أصل عشر، وسجّل 8.838 مقابل 6.862 في المتوسط — "واحدة من أكثر القفزات الذكائية إثارة للإعجاب" التي شاهدها المختبر. كما استخدم ما يقرب من ثلاثة أضعاف عدد الرموز، وكان أبطأ بشكل ملحوظ؛ إذ استغرقت إحدى المهام أطول بنحو 600%. لذا فإن قفزة الجودة حقيقية، وكذلك الثمن من حيث الوقت الفعلي.

Four more things to hold against it:

No third-party benchmarks yet. Every headline figure in this article is Ali​baba-reported as of August 15. The vendor card is detailed, but reproduction by an independent lab has not happened. If your decision depends on verified numbers, wait for them — the weights will still be there.

VRAM floor is real. BF16 needs an 80 GB-class GPU. To fit a 24 GB card you must run 4-bit, and community reports (dev.to comment thread, days after release) say quantized builds "lose focus after long context." Official weights if you have the VRAM; quantized if you do not.

1M context is hosted-only. The open weights cap at 262K. The extendable-to-1M figure is a Qwe​n Cloud hosted feature, not something a local copy does out of the box.

"Beats Opus" needs qualifiers. Agentic benchmarks reward harness-specific behaviors, and a top comment on the dev.to launch post put it bluntly: "they do not beat opus on real-world usage." Treat the scoreboard as an upper bound on a good day, not a promise.

How it sits in the Qwe​n 3.8 family

vs Qwen3.6-27B — the same hardware class and 262K context, but a major capability jump at the cost of speed and token efficiency. If you already run 3.6 and are throughput-bound, staying is defensible.

vs Qwen3-Coder-30B-A3B — the Coder is a MoE with roughly 3.3B active parameters that runs much faster (~90–110 tokens/s). The dense 27B is slower per token but inherits the generation's agentic gains; if the 27B's software-engineering scores hold up under independent testing, the Coder-30B's role shrinks.

vs Qwen3.8-Max — the 2.4T MoE flagship is a datacenter model (API-only, around $2/$6 per million tokens in published coverage) with 1M context and video. The 27B is the deployable one. They answer different questions.

Cost: the local-versus-API question, settled

For a closed model you compare per-token prices. For Qwen3.8 27B the marginal token costs nothing on your own hardware — the real price is the upfront GPU and your time. At 4-bit that is a 24 GB card; at BF16 it is an 80 GB-class machine. If you only need to evaluate the model, or you lack the hardware, the hosted route exists: OrcaRouter now lists Qwen3.8 27B with a free, rate-limited tier that bills $0 and returns HTTP 429 past its cap, plus a paid tier at $0.33 per million input and $2.40 per million output tokens (checked August 15). Qwe​n Cloud's hosted version, with the 1M context, is marked "coming soon."

Cost comparison card for Qwen3.8-27B titled 'Self-host vs hosted — the real price': Apache 2.0 weights at $0 license plus your own GPU and electricity; a 4-bit GGUF of roughly 17 GB that fits a 24 GB card; BF16 at 55.6 GB needing an 80 GB-class GPU; a free rate-limited hosted tier at $0 with HTTP 429 past the cap; and a paid hosted tier at $0.33 input / $2.40 output per million tokens with a 262K context window. Footer: prices from the OrcaRouter model page, read August 15, 2026.

The honest framing: for an open-weights model, a provider is a convenience, not a dependency. The free tier is the cheapest possible way to decide whether a 55.6 GB download is worth it. But once you have decided, the weights are the thing — and they are free.

How to actually try it

Fastest evaluation — hit the free rate-limited hosted tier; if it answers what you need, you are done and it cost nothing.

Real test on your hardware — download a community GGUF (the Q4_K_M build is roughly 17 GB and fits a 24 GB card) and run it in llama.cpp or Ollama. Pass the Jinja chat template or the model answers you in thought tags. For vision from a GGUF you also need the separate mmproj file. Our runbook on running Qwen3.8 27B locally walks through it.

Full precision — huggingface-cli download Qwen/Qwen3.8-27B onto an 80 GB-class GPU.

Verify what you downloaded — check the publisher is the Qwe​n org and validate shards against crc32.txt; lookalike repos appeared before release.

When this review is wrong (and who should skip Qwen3.8 27B)

You have no GPU and will not rent one. The free tier is rate-limited and the paid tier is $0.33/$2.40 per million — a small API model may serve you cheaper and more reliably. This model is for people who want to own the inference.

You need an SLA. The hosted tier is days old; the OrcaRouter model page, read today, shows a 33.3% error rate over the trailing seven days and p50 first-token latency of 225 ms. That is a brand-new model under early load, not a production contract. Self-hosters should pin their download commit and keep a fallback.

You need verified benchmarks for a purchase decision. Vendor-reported numbers are directionally useful, not procurement-grade. Wait for independent reproduction.

262K open-weights context is not enough. The 1M extension is hosted-only today.

Throughput is your bottleneck. Bulk summarization or large batches will hurt: this is a dense 27B that also burns roughly 3x the tokens of its predecessor. For cheap-bulk work, a smaller or faster model wins.

You are on a 12 GB card. 2-bit GGUFs squeeze in with visible quality loss; this is not the model for you.

Verdict infographic titled 'Use it if / Skip it if' for Qwen3.8-27B: left lane in soft blue labeled 'Use it if' with three items — '24 GB GPU', 'Free Apache 2.0 weights', '262K context + image/video'; right lane in soft gray labeled 'Skip it if' with three items — 'No GPU', 'Need an SLA', 'Need verified benchmarks'.

الخلاصة

Qwen3.8 27B is the strongest open-weights 27B available today, and it is free forever under Apache 2.0. If you have a 24 GB+ GPU and want agentic coding, 262K context, and native image/video input without a metered bill, this is the buy. The reasons to hold off are equally concrete: you need independent benchmark confirmation, an SLA, more than 262K open-weights context, or higher throughput than a dense 27B gives you. The model is out, the weights are real, and the numbers are good — just remember whose numbers they are.

© 2026 OrcaRouter

لمقدمي الخدمات

هل تدير منصة استدلال؟ اعرض نماذجك على OrcaRouter.

تواصل معنا

انضم إلى مجتمعنا

DiscordEmailXGitHubYouTube