GPT-5.6 인하 후 가격: Luna vs Terra vs Sol, 어떤 등급을 사용해야 할까?
Guides & Insights

GPT-5.6 인하 후 가격: Luna vs Terra vs Sol, 어떤 등급을 사용해야 할까?

작성자

Jim Song

게시일

최신 모델 · 20모든 모델 보기
벤치마크: Artificial Analysis · 매일 업데이트
모든 게시물로 돌아가기

After O​penAI's July 30, 2026 price cut, the GPT-5.​6 family — GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol — has never been cheaper to run at scale. A week later, on August 6, O​penAI pushed the same tiers into ChatGPT itself: GPT-5.6 Luna is becoming the default model for free and Go accounts with unlimited text chats, while GPT-5.6 Sol was updated for Plus and Pro. But the three tiers are priced very differently, and picking the right one (plus using caching, batch, and routing well) is the single biggest lever on your API bill. This guide lays out the full post-cut pricing, worked cost examples, how each tier compares to rivals, and how to route between them through one endpoint.

Prices are per million tokens (input / output), reflecting the post-cut rate card as reported and OrcaRouter's pass-through model pages; tiered and cache rates come from OrcaRouter's model pages, and competitor figures are attributed. The ChatGPT rollout details in this piece are as O​penAI stated them on August 6, 2026 — vendor-reported, not independently measured. Prices change — verify before building.

TL;DR

GPT-5.6 Luna ($0.20 / $1.20) is the high-volume workhorse; Terra ($2 / $12) is the balanced middle; Sol ($5 / $30) is the flagship for the hardest reasoning and agentic coding. All three share a ~1.05M-token context and 128K max output. Prompt caching and the batch option cut costs further, while very long inputs move you to a higher tier. On the consumer side, O​penAI announced August 6 that GPT-5.6 Luna is becoming the default for Free and Go ChatGPT accounts with unlimited text-only chats, and that an updated GPT-5.6 Sol is rolling out to Plus and Pro. Use Luna for the easy majority, Terra when you need more, and Sol only where it earns its cost — and route by difficulty through one O​penAI-compatible endpoint to minimize spend.

주요 시사점

• Luna: 1M 토큰당 $0.20 / $1.20 — 빠르고 저렴하며, 대용량 및 지연 시간에 민감한 작업에 적합합니다.

• Terra: $2 / $12 — 플래그십이 필요 없는 더 어려운 작업을 위한 균형 잡힌 중간 티어

• Sol: $5 / $30 — 심층 추론, 대규모 코딩, 장기 지평 에이전트를 위한 플래그십 모델.

• 캐싱과 배치 처리는 비용을 더욱 절감하며, 긴 입력의 경우 더 비싼 긴 컨텍스트 티어로 이동하게 됩니다.

• ChatGPT (O​penAI-stated, Aug 6): Luna becomes the free/Go default with unlimited text chats; Sol gets a reliability update for Plus/Pro.

• 가장 큰 비용 절감 수단: 난이도별 라우팅 — Luna가 처리할 수 있는 작업에 Sol 요금을 지불하지 마세요.

The new GPT-5.​6 base pricing

Here's the post-cut base pricing per million input/output tokens: Luna $0.20 / $1.20 (down from $1 / $6), Terra $2 / $12 (down from $2.50 / $15), and Sol $5 / $30 (unchanged). The two cheaper tiers were cut on July 30, 2026; the flagship held. Put simply, the low tiers are now priced for scale, while the flagship stays premium.

계층형, 캐시형, 배치형 가격 책정

The base rate is only the starting point. GPT-5.​6 pricing has three modifiers that materially change your effective cost:

• 긴 컨텍스트 등급. 매우 긴 입력에 대해 가격이 단계적으로 올라갑니다. OrcaRouter의 패스스루 페이지에서 Luna는 대형 컨텍스트 등급까지 $0.20 / $1.20이고, 그 이상에서는 $0.40 / $1.80입니다. Sol은 기본 등급에서 $5 / $30, 최대 등급에서는 $10 / $45입니다. 등급은 각 요청의 입력 토큰 수에 따라 선택되므로, 몇 개의 큰 프롬프트가 조용히 평균 가격을 올릴 수 있습니다.

• 프롬프트 캐싱. 재사용된 컨텍스트는 큰 할인율로 청구됩니다 — Luna의 캐시 읽기는 대략 토큰 100만 개당 $0.02(신규 $0.20 대비)이며, 캐시 쓰기는 약 $0.25입니다. Sol의 캐시 읽기는 약 $0.50(신규 $5 대비)입니다. 시스템 프롬프트가 안정적인 에이전트와 채팅의 경우, 캐싱은 종종 가장 큰 절감 효과를 제공합니다.

• 배치(Batch). 긴급하지 않은 작업은 비동기적으로 할인된 가격으로 실행됩니다 — 대량 분류, 평가(evals), 지연 시간이 중요하지 않은 오프라인 생성에 이상적입니다.

계산 예시: 실제 비용은 얼마인가

구체적인 숫자가 티어의 실제를 보여줍니다. 월 1,000만 토큰을 70% 입력 / 30% 출력 비율로 사용하는 경우를 예로 들어 보겠습니다(일반적인 채팅/에이전트 혼합 비율).

• Luna: 한 달에 약 $5 (프롬프트 캐싱 사용 시 대략 $4.37).

• Terra: 월 약 $50.

• Sol: 월 약 $125 (캐싱 적용 시 약 $109).

같은 볼륨에 대해 Luna와 Sol 사이의 25배 차이가 있습니다. 이것이 바로 각 요청을 가장 저렴하면서도 적합한 티어에 매칭하는 것이 어떤 단일 요금보다 더 중요한 이유입니다. (이 수치는 OrcaRouter의 페이지 내 비용 계산기와 일치하며, 이 계산기는 표시 가격을 기준으로 추정합니다. 실제 수치는 캐싱 및 입출력 구성에 따라 달라집니다.)

각 티어가 실제로 무엇을 위한 것인지

GPT-5.6 Luna — 대용량 처리의 일꾼

Luna is the fast, cost-efficient tier, tuned for high-volume, latency-sensitive workloads: chat, classification, extraction, routing, and lightweight agentic tasks, with a p50 time-to-first-token around 1.65 seconds. After the cut, at $0.20 / $1.20 it's priced to compete directly with cheap open models — the default for the easy majority of calls. It is also the tier O​penAI is steering consumer ChatGPT toward: per its August 6 announcement, GPT-5.6 Luna is becoming the default model for Free and Go accounts, with text-only chats going unlimited and a new Think button for higher reasoning on harder questions arriving the following week (file, image, and voice limits remain). O​penAI says Luna makes 62% fewer factual errors than the GPT-5.5 Instant it replaces — a vendor-reported figure, but a clear signal that the volume tier is where the company is pointing most of its traffic.

GPT-5.6 Terra — 균형 잡힌 중간

Terra는 볼륨(volume)과 플래그십(flagship) 사이에 위치합니다. 더 까다로운 추론과 코딩에서는 Luna보다 뛰어나지만, Sol보다는 훨씬 저렴하죠. $2/$12라는 가격으로, Luna로는 부족하지만 플래그십까지는 필요 없는 경우 — 중간 복잡도의 추출, 초안 작성, 그리고 대규모로 실행되는 다단계 작업에 적합한 합리적인 기본 선택입니다.

GPT-5.6 Sol — 플래그십

Sol is built for the hardest work: deep multi-step reasoning, large-scale software engineering, and long-horizon agentic workflows, staying coherent across a ~1.05M-token context and up to 128K output. At $5 / $30 (base) it's a premium choice — reserve it for tasks that genuinely need it, like complex multi-file coding or long agent runs. On August 6, O​penAI also updated GPT-5.6 Sol in ChatGPT for Plus and Pro users: the company says the chat version is more reliable with facts — 68% fewer factual errors, per O​penAI — and gives more focused answers, with a new slider to control how much reasoning effort it applies. The update is chat-only (the Sol behind Work and Codex is unchanged), and the API tier stays at $5 / $30.

추론 토큰 주의사항

GPT-5.​6 are reasoning models, so effective output cost can exceed a naive estimate: when reasoning is on, internal reasoning tokens are billed as output. On hard prompts with high reasoning effort, that can dominate your bill. Tune reasoning effort to the task (low or off for simple calls), cap output tokens where you can, and measure actual usage. Cheaper per-token rates help; generating fewer tokens helps more.

각 등급이 경쟁사와 어떻게 비교되는지

The cut repositioned GPT-5.​6 against the field. Per reporting, Luna's $0.20 / $1.20 now undercuts Claude Haiku 4.5 by roughly 5x on input and 4x on output, and Terra's $2 / $12 falls below Claude Sonnet 5's standard pricing (reported at $3 / $15). Against open weights, Luna sits near the floor set by Deep​Seek's V4 line and Zhipu's GLM-5.2 (about $1.20 / $4.10), with Qw​en and Gemini Flash tiers nearby. The upshot: for cost-sensitive work, Luna is now competitive with the cheapest capable models rather than a premium alternative to them — while Sol remains a genuine premium tier for capability you can't get cheaply.

진정한 절감 수단: 난이도별 라우팅

The cheapest bill isn't a single tier — it's matching each request to the least expensive model that can do it. In practice: send easy, high-volume calls to Luna (or a cheap open model), step up to Terra for harder tasks, and use Sol only for the genuinely difficult minority. Combined with caching and batch, this routinely cuts costs far more than any single price change. The catch is operational: you don't want to re-integrate three O​penAI tiers plus open-model alternatives separately.

세 가지 모두(및 더 저렴한 경쟁 제품)를 하나의 엔드포인트로 이용하세요.

This is where a vendor-neutral router helps. OrcaRouter exposes GPT-5.6 Luna, Terra, and Sol — at the same post-cut prices, 0% markup — through one O​penAI-compatible endpoint, alongside cheaper open models like DeepSeek V4 Pro, GLM-5.2, and Qw​en. Switching tiers (or A/B testing Luna against an open model) is a config change, not a re-integration, and you can route each request to whatever is cheapest.

The post-cut Luna rate card is visible on OrcaRouter's own model page at list price — $0.20 / $1.20 per million tokens — because the router passes provider prices through at 0% markup, with an on-page cost calculator for a typical monthly bill. There's also a free tier and an Offers page for additional savings.

자주 묻는 질문

What are the GPT-5.​6 prices after the cut?

Luna는 입력/출력 토큰 100만 개당 $0.20 / $1.20, Terra는 $2 / $12, Sol은 $5 / $30입니다. Luna와 Terra는 2026년 7월 30일에 인하되었으며, Sol은 변경되지 않았습니다.

Which GPT-5.​6 tier should I use?

대용량·저지연 작업에는 Luna, 플래그십이 필요 없는 더 까다로운 작업에는 Terra, 가장 어려운 추론 및 에이전틱 코딩에는 Sol. 난이도에 따라 라우팅하여 비용을 최소화하세요.

Is GPT-5.6 Luna free in ChatGPT?

O​penAI announced on August 6 that GPT-5.6 Luna is becoming the default model for free and Go ChatGPT accounts, with unlimited text-only chats; limits remain on file uploads, images, and voice tools, and a Think button for higher reasoning is rolling out separately. These are O​penAI-stated plans, not independently verified.

How much does GPT-5.​6 cost per month?

월 1천만 토큰 기준(입력 70%): 캐싱 전 기준으로 Luna에서는 약 $5, Terra에서는 약 $50, Sol에서는 약 $125 — 25배 차이입니다. 실제 비용은 캐싱과 입력/출력 비율에 따라 달라집니다.

세 가지 등급 모두 동일한 컨텍스트 창을 가지고 있나요?

네 — 대략 1.05M 토큰의 컨텍스트와 최대 128K 출력 토큰을 지원하며, 비전, 도구, JSON 및 추론 기능을 제공합니다.

프롬프트 캐싱은 비용에 어떤 영향을 미치나요?

캐싱은 반복 컨텍스트 비용을 크게 줄여줍니다 — Luna의 캐시 읽기는 토큰 100만 개당 약 $0.02인 반면, 새로 읽으면 $0.20입니다 — 따라서 가능하면 캐시된 프롬프트를 재사용하세요. 반대로 긴 입력은 더 비싼 장문 컨텍스트 등급으로 이동시킬 수 있습니다.

Anthropic이나 오픈 모델과 비교했을 때 티어는 어떤가요?

Per reporting, Luna undercuts Claude Haiku 4.5 (~5x input / 4x output) and Terra falls below Claude Sonnet 5 ($3 / $15). Luna is also near the cheap open-model floor (Deep​Seek, GLM-5.2 ~$1.20/$4.10).

세 가지 등급을 모두 함께 어디에서 사용할 수 있나요?

Through OrcaRouter's single O​penAI-compatible endpoint at 0% markup, alongside cheaper open models for difficulty-based routing.

결론

After the cut, GPT-5.6 Luna ($0.20 / $1.20) and Terra ($2 / $12) are compelling for high-volume and mid-tier work, while Sol ($5 / $30) remains the premium flagship — a 25x cost spread that rewards smart routing. The consumer rollout reinforces the split: O​penAI is making Luna the free default while reserving Sol's reliability update for paid plans, which for API builders is one more reason to route easy traffic to Luna. The biggest savings come from routing by difficulty, leaning on caching and batch, and controlling reasoning tokens rather than defaulting to one tier. Doing that is easiest through a single 0%-markup endpoint like OrcaRouter, where all three tiers (at the new prices) sit next to the cheaper open models you'll want to compare them against.

이 글에서 비교한 모델2

이 글에서 자동 인식 · 벤치마크: Artificial Analysis · 매일 업데이트

© 2026 OrcaRouter

제공업체용

추론 플랫폼을 운영하시나요? OrcaRouter에 모델을 등록하세요.

문의하기

커뮤니티에 참여하세요

DiscordEmailXGitHubYouTube