AI News | General
Anthropic’s Claude Opus 5.5 Challenges OpenAI’s GPT-6 Sol and GPT-6 Luna
AI got cheaper this week, but the tests tell a more complicated story. OpenAI and Anthropic released new models whose value depends on the task. Alibaba unveiled a chip aimed at powering its next generation of models, while Anthropic’s AI helped uncover a biological system scientists still cannot explain. Governments, meanwhile, are asking who should test the most powerful systems before they reach the public.
Sep 24th, 2026
klarapelhe.rs/en/ai-news
Anthropic’s Claude Opus 5.5 Challenges OpenAI’s GPT-6 Sol and GPT-6 Luna
Key takeaways
- OpenAI halved API token prices for its new GPT-6 Sol and Luna models; independent tests found substantial cost savings but mixed capability changes.
- Anthropic released Claude Opus 5.5 with lower prices and stronger scores in independent evaluations; METR found a modest, rather than dramatic, advance in AI research capability.
- Alibaba unveiled a new accelerator and a plan for much larger Qwen models, while Anthropic reported an AI-assisted biological finding whose function remains unknown.
- Leaders from Europe and several other regions called for independent frontier-model testing and explored a global oversight institution; their statement creates no binding rules.
OpenAI realeases GPT-6 Sol and Luna and significantly cuts their costs
OpenAI released GPT-6 Sol and Luna on Tuesday as lower-cost members of the GPT-6 family introduced with Astra earlier this month. In the API, Sol costs $2 per million input tokens and $10 per million output tokens, while Luna costs $0.10 and $0.50. Those are roughly half the previous GPT-5.6 Sol and Luna rates. OpenAI says the models use training methods related to Astra's and bring improvements in coding, computer use and factuality to less expensive tiers. Sol and Luna are available through the API and in ChatGPT Work and Codex for paid plans; free and Go users receive Luna in the desktop app, with rollout staged through the day of release.¹
For builders, the price change is straightforward; the performance change depends on the job. Artificial Analysis ran the models in its own evaluation setup and found that the overall Intelligence Index and Coding Agent Index scores remained broadly level with the GPT-5.6 predecessors. Sol at maximum effort gained two points on its Coding Agent Index, while Luna lost two. Its measured cost per Intelligence Index task fell from $1.99 to $1.06 for Sol and from $0.18 to $0.07 for Luna. These are costs under a particular benchmark workload, not a guarantee for an application's bills.²
The independent testing also exposed a tradeoff hidden by a single headline score. In Artificial Analysis's knowledge test, Sol produced fewer hallucinations partly because it declined to answer more questions: its attempt rate fell from 99% to 83%, and accuracy on the test fell five points. Both new models also regressed on its professional-work evaluation, where shorter deliverables omitted required elements more often. OpenAI's own announcement highlights stronger scores on several agent and coding tests, but its evaluations may use different prompts, tools and effort settings from production deployments.¹,²
A team processing large volumes of well-defined tasks can now test the same workflow at a lower token price. A team buying completeness or sustained tool use should compare actual outputs, failure rates and total cost per successful task before switching. The practical story of this release is therefore a lower cost floor with uneven quality changes across tasks, not a uniform leap in model ability.²
Claude Opus 5.5 is now live and pairs lower prices with stronger independent results
Anthropic launched Claude Opus 5.5 same day when OpenAI has released its models, and this the first model in its 5.5 family, at $4 per million input tokens and $20 per million output tokens, down from $5 and $25 for Opus 5. Cache reads fell from $0.50 to $0.20 per million tokens. Anthropic says lower token use on typical workloads makes total task cost about 40% lower than Opus 5 and that output is more than 30% faster. It also raised five-hour usage limits on several subscription plans. These claims refer to Anthropic's workloads and pricing; a customer's mix of cached input, generated output and retries will determine its own savings.³
Independent measurements support a substantive advance, while putting boundaries around it. Artificial Analysis placed Opus 5.5 first on its Intelligence Index at maximum effort, with a score of 58, and found leading results on six of its ten component evaluations. On its professional-work AA-Briefcase test, the model reached 1,822 Elo, 143 above Claude Fable 5.1. Yet the same evaluator measured roughly 119,000 output tokens per task at maximum effort, far more than Opus 5's 73,000, so the per-token discount need not translate into the same saving for every task.⁴
Anthropic emphasizes long coding sessions and says Opus 5.5 includes action screening, an auditable sandbox and stronger resistance to prompt injection. In its own HAProxy code-translation exercise, the new model finished faster and for less money than Fable 5.1, with both rewrites passing nearly all of the project's regression tests. Those are vendor-run examples, not evidence that unattended changes are safe in every codebase.³
On the other hand, METR, which assessed the model before release, reported a narrower conclusion about its ability to accelerate AI research. It found a likely modest improvement over Fable 5.1, but no evidence of a large jump in the judgment needed for fully autonomous research. The group also noted weaknesses on hard, long tasks. METR conducted the assessment under an unpaid agreement; Anthropic had a chance to review and edit the summary, and METR signed off on the final text.⁵ For organizations using agents, Opus 5.5 is a strong candidate for demanding work, but approval should rest on tests of the work they actually delegate and on the safeguards around irreversible actions.
Alibaba unveils a new AI chip while training Qwen 4
At its Apsara conference in Hangzhou two days ago, Alibaba disclosed that Qwen 4 is in training and that later Qwen 4.5 and Qwen 5 models are planned at five to ten trillion parameters. The company also unveiled the Zhenwu V900, an accelerator designed by its T-Head unit for training and inference, the computation used to run a trained model. Alibaba says it has three times the performance of the M890 introduced in May, 216 GB of memory and 1,200 GB/s of inter-chip bandwidth. Commercial release and mass production are scheduled for the first quarter of 2027. These specifications and comparisons are company claims; the announced chip is not yet a generally available product with independently established performance.⁶
The announcement ties a model roadmap to a domestic compute stack. Alibaba says its existing Zhenwu chips serve more than 650 customers and that its new supernode design can support clusters of up to 500,000 cards. It also set a target for Alibaba Cloud to operate more than 20 gigawatts of global data-center capacity by 2032. Neither the planned model size nor the capacity target establishes that the models or sites will be delivered on schedule. Parameter count is a measure of model scale, not a direct measure of useful performance.⁶
Alibaba used the conference to describe an experiment in which Qwen3.8-Max completed 33 automated cycles of training optimization over a month, improving its Artificial Analysis score from 40 to 45. That is a reported internal process and outcome; the public announcement does not provide enough detail to establish whether the gain generalizes beyond the cited measure. Associated Press reporting placed the launch amid China's effort to build alternatives to Nvidia hardware under US export restrictions. Analysts it interviewed cautioned that chip design alone cannot close a computing gap without manufacturing capacity and other supply-chain advances.⁶,⁷
Anthropic's Claude makes a significant scientific breakthrough
Anthropic reported that Claude agents searching DNA-sequence data identified an unusual system in bacteriophages, viruses that infect bacteria. The finding came from a new internal life-sciences laboratory combining computational search with human-run experiments. Anthropic calls the system array-associated reverse transcriptases, or ARTs, after the enzyme family and nearby repeated DNA sequences. The underlying reverse transcriptase was already known; Anthropic says the previously unnoticed combination of the enzyme, an accessory protein and a repeat array was the discovery.⁸
The company says about 950 agents processed 210 million tokens over 21 hours. They surveyed more than 200,000 reverse transcriptases, surfaced roughly 3,500 candidate systems and narrowed the field to 20 reports for closer analysis. An agent noticed a repeat array beside an unusual enzyme gene. Anthropic's scientists then tested the candidate in the lab and found that the array is expressed as distinct short RNA molecules. Human researchers performed the physical experiments; the agents generated and investigated hypotheses.⁸
Anthropic compares the arrangement with CRISPR because both feature repeated sequences and associated molecular machinery. That structural resemblance does not establish the ART system's function. The company explicitly says it does not yet know what ARTs do, and The Verge notes that no practical application has been shown. Anthropic says it released a preprint, but the linked technical report was not accessible during this review, so the account of methods and experiments here rests on its public description rather than inspection of the preprint.⁸,⁹
International leaders ask for independent testing of frontier AI models
A statement led by Finland and Norway and published on the Dutch government's site on 22 September asks AI companies to adopt transparent safety protocols, including testing before deployment and independent evaluation with sufficient evaluator access. It asks governments to coordinate standards and serious-incident reporting, and invites UN member states to explore an institution that could set standards and enable verification when capability thresholds are crossed. The declaration says frontier AI should remain under human direction and international law.¹⁰,¹¹
The signatories span Europe, Asia, Africa, the Middle East and North America; the published list includes leaders from the European Commission, Germany, Canada, Kenya, Singapore, South Africa and the United Arab Emirates. The Finnish AI Region reported that the declaration was adopted alongside the UN General Assembly on 21 September and that neither the United States nor China signed it. The text itself remains open to further endorsements.¹⁰,¹¹
For model developers, the proposed direction would make external access, pre-release evaluation and incident reporting more central to release decisions if governments later implement it. For smaller countries, the statement also calls for access to scientific capacity and trusted evaluation, recognizing that oversight cannot work if only the largest AI powers can inspect frontier systems. These are proposals, not operational requirements: the statement sets no binding testing threshold, deadline, enforcement mechanism or institution.¹⁰
Sources
- OpenAI. Introducing GPT-6 Sol and Luna. 22 September 2026.
- Artificial Analysis. GPT-6 Sol and Luna push the cost efficiency frontier. 22 September 2026.
- Anthropic. Claude Opus 5.5. 22 September 2026.
- Artificial Analysis. Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index. 22 September 2026.
- METR. Summary of METR's predeployment evaluation of Claude Opus 5.5. 22 September 2026.
- Alibaba Cloud. Alibaba Unveils Roadmap on Full-Stack AI Strategy from Chips, Cloud Infrastructure, Models to Agents. 22 September 2026.
- Chan Ho-Him. China's Alibaba unveils new powerful chip and ambitious AI model plans. 22 September 2026. Associated Press.
- Anthropic. Claude discovers a novel enzyme system with CRISPR-like repeats. 23 September 2026.
- Robert Hart. Anthropic's biolab made a discovery it's comparing to Crispr. 23 September 2026. The Verge.
- Government of the Netherlands. A Call for Control of Frontier AI Models. 22 September 2026.
- Martti Asikainen. Finland and Norway lead 22-nation call for global AI oversight body, as US and China stay out. 22 September 2026. Finnish AI Region.