AI Tools / AI news for Malaysia
OpenAI cut GPT-5.6 Luna prices by 80%. What Malaysian teams should count
The headline token rate is much lower, but a useful AI workflow still includes retries, tools, storage, oversight and exchange-rate movement.

In brief
- OpenAI says GPT-5.6 Luna now costs US$0.20 per million input tokens and US$1.20 per million output tokens after an 80% reduction.[1]
- GPT-5.6 Terra now costs US$2 per million input tokens and US$12 per million output tokens after a 20% reduction.[1]
- OpenAI says Sol Fast can deliver up to 2.5 times standard speed at twice the price, with no change in intelligence.[1]
Luna became cheaper, while Sol added a speed premium
OpenAI's late-July pricing move made GPT-5.6 Luna unusually cheap for large volumes of everyday AI work. One million input tokens now costs twenty US cents, while one million output tokens costs US$1.20.[1]
For Malaysian teams, the ringgit amount changes with the exchange rate and payment charges, so this article does not freeze a possibly stale conversion. More importantly, the invoice also depends on prompt size, generated output, repeated attempts, tools and the amount of human checking needed to get a usable result.[1]

What the new Luna rate changes
GPT-5.6 Luna now has a published price of US$0.20 per million input tokens, US$0.02 for cached input and US$1.20 per million output tokens. OpenAI describes the input and output change as an 80% reduction.[1]
That can materially improve the economics of high-volume work such as classification, extraction, drafting and internal assistants. The saving is largest when a team already understands its traffic and can route routine tasks to Luna without increasing the failure or retry rate.[1]
OPENAI / GPT-5.6 LUNA
The unit price fell 80%. Usage discipline still decides the bill.
OpenAI states an 80% reduction. The earlier input and output rates below are mathematical implications of that statement, not separately quoted historical prices.[1]
By Utopia Data & AI Team
Input / 1M tokens
- Implied before
- US$1.00
- Now
- US$0.20
Output / 1M tokens
- Implied before
- US$6.00
- Now
- US$1.20
Cached input / 1M
- Implied before
- Not stated
- Now
- US$0.02
80% LOWER: OPENAI'S STATED CHANGE
Why token price is not the total cost
A production workflow may call search, databases or other tools, store files and logs, retry an uncertain answer and ask a person to review the result. Those costs sit outside the headline token rate, yet they determine whether the workflow saves money after it reaches real users.[1]
Long context also needs attention. OpenAI says prompts above 272,000 input tokens are priced at twice the input rate and one-and-a-half times the output rate for the full request. Uploading everything by default can therefore erase part of the price advantage.[1]

How Malaysian teams should test the saving
The cleanest comparison is cost per accepted outcome in ringgit. Teams should record tokens, tool calls, retries, review time and success for the same task before and after a routing change, then use the exchange rate and charges from the actual invoice.[1]
A cheaper model can still be expensive when it produces more unusable outputs. Conversely, Luna may be the better business choice even when a stronger model scores higher on a benchmark, provided it meets the acceptance threshold for a well-defined routine task.[1]

Why Malaysia should care
Lower input and output prices make high-volume assistants, extraction and classification more practical for Malaysian teams. Budgeting still needs a live USD-to-MYR rate and a measurement of the entire completed workflow.
Malaysian SMEs
More routine AI tasks may now clear a cost-benefit threshold at higher volumes.[1]
Practical move: Run one month of real traffic and divide the total bill by successful completed outcomes.
Developers
Model routing can use Luna for simpler work and reserve stronger or faster options for difficult steps.[1]
Practical move: Measure quality and retry rates before routing everything to the lowest token price.
Finance teams
The supplier rate is in US dollars and usage is variable.[1]
Practical move: Set a ringgit budget buffer, usage alerts and an owner for unexpected growth.
What Malaysians can do now
- 01Record input tokens, output tokens, retries and successful outcomes for each production workflow.
- 02Convert the live US-dollar invoice into ringgit using the actual card or billing rate, not a static article estimate.
- 03Test a cheaper model against your acceptance criteria before changing production routing.
What we still do not know
A lower list price does not reveal your final monthly bill
- The exchange rate and payment charges each Malaysian customer will incur.
- How much model quality, retries and human review differ for each real workflow.
- Whether these rates or available speed modes will change again after this review date.
Sources
- 1.Building abundant intelligence OpenAI, 31 July 2026
- 2.GPT-5.6 Luna Model | OpenAI API OpenAI


