Tech News, powered by Quinck: three stories from this week that seem to belong to different worlds. A chip from OpenAI, a model from Anthropic criticised by its own users, two Apple processors. In reality they describe the same shift. Raw computing power has stopped being the bottleneck. Now the problem is how much it costs to deliver intelligence, and how much it costs to receive it.
The 3 stories in 30 seconds
- OpenAI has published the first numbers for Jalapeño, the inference chip designed with Broadcom: up to 1.9 times more useful work per kilowatt and up to 3.6 times lower end-to-end latency compared with the reference systems. In-house measurements, on engineering silicon, against Blackwell: not against Rubin.
- Claude Opus 5 sits at the top of the intelligence rankings and at the bottom of the satisfaction ones. The reason is not the quality of its answers: it is their length. Anthropic had documented the behaviour before launch, and on 20 August it released a "Concise" output style to contain it.
- Apple has announced M6, its first 2-nanometre chip, and M5 Ultra, the most powerful ever produced in Cupertino. The Pro, Max and Ultra versions of the M6 are missing: according to Bloomberg they are not late, they are cancelled.
1. Jalapeño: OpenAI brings the numbers, and the comparison is with Blackwell
On 24 June OpenAI and Broadcom had presented Jalapeño, the first accelerator designed by OpenAI: a reticle-sized ASIC that reached tape-out in nine months, a timeline that is almost anomalous in the world of custom silicon. But the announcement was missing the part that matters: the data. Hence the scepticism with which the news was received in June.
On 25 August the data arrived. OpenAI published the first measurements on working silicon, run with InferenceX, the SemiAnalysis benchmark that evaluates the full path of a request instead of a synthetic kernel. The workloads are three open models: GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 with a trillion parameters.
The declared results:
| Metric | Jalapeño vs reference systems |
|---|---|
| Peak AI work per kilowatt | from 1.5× to 1.9× |
| End-to-end latency | from 1.7× to 3.6× lower |
| Package TDP | 700 W (versus 1,200 W and 1,400 W) |
The interesting technical point is not any single number, but the combination. In inference systems, throughput and latency are bought at each other's expense: you increase the batches, the cost per token goes down, the response time goes up. Jalapeño claims to improve both with the same architecture, having optimised in particular the prefill and communication phases, that is, the points where data moves instead of being computed. It is exactly the bottleneck that agents make more painful: an agent chains many steps, and every delay adds up across the whole task.
Two caveats that change the reading
The comparison is with Blackwell, not with Rubin. The reference systems are the Nvidia platforms currently in the field; for the Rubin generation, which is the real point of comparison in time for a chip being deployed at the end of 2026, there are no numbers. It is the question that remains open, and it will decide whether Jalapeño is a structural advantage or a snapshot of a moment.
The numbers are OpenAI's. The hardware is not available to outside teams, nobody has independently re-run the tests, and efficiency is normalised on each accelerator's declared package TDP, which is a methodologically legitimate choice but not a neutral one. The workloads are nominal: 8,000 input tokens, 1,000 output tokens, single-token prediction.
What remains are solid facts: the chip exists, it runs, and the results are operational, not theoretical peaks. Internal deployment starts by the end of 2026, in small volumes, with real scale expected in 2027; generations 2 and 3 are already in the works. Broadcom's Hock Tan spoke of operating costs around 50% lower than the GPUs commonly used for inference.
For anyone buying inference (which, by now, means anyone building software) the consequence is not immediate: Jalapeño is not something you buy, you will use it by serving OpenAI models. But the signal is that the cost per token of the big providers still has a lot of room to fall, and that the room will come from vertical silicon, not from commercial discounts.
2. Opus 5: first in the rankings, last in readability
Claude Opus 5 came out on 24 July, positioned by Anthropic as close to the frontier intelligence of Fable 5 at half the price. On benchmarks the promise holds: the model sits at the top of the aggregate intelligence rankings and competes with GPT-5.6 on coding, agentic search and computer use.
Then the community arrived, and the verdict was far less kind.
The problem is not the correctness of the answers. It is that Opus 5 produces roughly twice as many output tokens as previous models, with a tendency to explain, preface, verify and delegate even when the question was trivial. On Reddit the joke that did the rounds had someone asking which ice cream flavour to pick and getting a business plan in return. Reviewers with early access used words like "neurotic" and "slop"; tool developers reported a tendency to treat every minor problem as a serious incident to be solved with hundreds of lines of code.
The most interesting detail is that the judgement splits along a precise line: people like the model when it works alone and dislike it when it works with them. Left for hours on a verifiable goal, it produces results that even its critics call excellent. Put into a conversation, it floods you.
What Anthropic did
It must be said that the behaviour is not a surprise that emerged afterwards: the Opus 5 prompting guide published by Anthropic already documented at launch longer-than-average responses, more narration during agentic tasks, spontaneous self-verification and a more casual delegation to subagents, with the explicit warning that on small tasks that delegation multiplies cost and time.
Documenting a behaviour, however, is not the same as choosing it together with the user. For weeks the public response stayed at the documentation, while complaints piled up; then came the direct acknowledgement from an Anthropic engineer and, on 20 August, a "Concise" output style built into Claude Code, which can be enabled from /config or via "outputStyle": "Concise" in the settings. In the meantime it became clear that the most common remedy, lowering the effort, does not work the way you hope: it reduces how much the model thinks, not how much it writes. The opposite works better, namely removing the old instructions like "double-check your work" from your prompts, which on this generation are redundant and feed exactly the behaviour you want to avoid.
Why this news is not a complaint
Here is a fact that is worth more than any single release, and that tends to be overlooked.
In a world where AI becomes more intelligent, more autonomous and able to handle an ever wider variety of tasks, the human remains the point where work is approved, corrected and put into production. If we want to keep control and collect the benefits of automation, we have to interface with these systems. And an opaque answer, or one that has to be decoded, breaks exactly the point of contact where control is exercised.
In economic terms: a model that takes three thousand words to say what needed forty has not got the content wrong, it has shifted the cost. It has moved it from computation to reading, that is, to the side where the bottleneck is human and does not scale. The synergy that the expected productivity gain should come from (the person who reads, verifies, corrects and decides) jams right there.
Watts per token on one side, words per answer on the other: they are the same efficiency metric applied to the two ends of the chain. At the first end the industry measures obsessively. At the second it has only just started.
(Yes, we are aware of the irony of writing three paragraphs about verbosity. We'll stop here.)
3. Apple: 2 nanometres on the desktop, and a generation that won't arrive
On 25 August Apple announced M6 and M5 Ultra, with a press release and no keynote: unusual in itself for a leap in manufacturing process.
M6 is the first Apple chip built on 2 nanometres (TSMC's N2 process): a 12-core CPU, a 12-core GPU with Neural Accelerator, a dual 16-core Neural Engine and up to 170 GB/s of unified memory bandwidth. It goes into the new Mac mini, where Apple claims up to 4 times the AI performance, graphics and storage up to 2 times faster and a CPU 40% faster than the previous generation. It is the chip for standard workflows.
M5 Ultra is another category. It is Apple's first quad-die chip: two dual-die M5 Max fused with the new generation of UltraFusion, reaching up to 36 CPU cores, 80 GPU cores, 1.2 TB/s of unified memory bandwidth (50% more than the M3 Ultra) and up to 512 GB of unified memory. It is only in the new Mac Studio, on sale from 22 September, with the 512 GB configurations expected at the end of October.
That number, 512 GB of unified memory accessible from the GPU, is the reason this announcement matters even to people who don't buy Macs. It means running locally models that today require a rented cluster, with the data never leaving the machine. It is the same data sovereignty question we discussed regarding open weights models, approached from the hardware side instead of the licence side.
The hole in the lineup
Something is missing, and it is glaring: there is no M6 Pro, M6 Max or M6 Ultra. The Mac mini pairs the M6 with an M5 Pro, and the Mac Studio ships with M5 Max and M5 Ultra. The whole high end stays on the previous generation.
The benign hypothesis is that Apple has held back the announcements for later, partly so as not to overshadow the M5 Ultra. The rumours say the opposite. According to Bloomberg's Mark Gurman, those variants are not postponed: they will not be developed, for the first time in the Apple Silicon era. The reported roadmap is base M7 in the first half of 2027, M7 Pro and M7 Max at the end of 2027, M7 Ultra in 2028: the latter described as a "dramatic" leap in AI performance and as a possible engine for Apple Intelligence servers from 2029.
The stated reason is just one: AI. Apple had planned major neural processing upgrades for the M7 family and decided it was worth accelerating the next generation instead of completing the current one.
It is a choice that says more than many press releases. Skipping half a generation of chips (the most profitable half, at that) to arrive sooner with the unit that runs the models means that AI has stopped being a feature the silicon must support and has become the criterion by which silicon is designed and scheduled.
The common thread: intelligence is abundant, efficiency is not
Three stories, one movement.
OpenAI designs its own silicon to lower the energy cost and latency of inference. Apple rewrites its own roadmap to bring model execution to the desk sooner. And Anthropic discovers that a model at the top of the intelligence rankings can end up frustrating because it spends badly the one resource that no manufacturing process can double: the attention of whoever is reading.
None of the three is about making models better. All three are about the cost of delivering that quality: in watts, in milliseconds, in words.
For those building software there are three operational points, and they are concrete.
Measure end-to-end latency, not the peak. It is the metric the user perceives and the one that accumulates along an agentic flow. Excellent peak throughput with a long tail gives a slow product.
Treat verbosity and output format as project parameters. They are not an aesthetic detail: they are cost per token, review time and the risk that an important instruction ends up buried. They need to be set, measured and verified like any other requirement, and it is worth cleaning up inherited prompts, because with this generation of models the old defensive instructions often make the result worse.
Keep the choice of model and supplier reversible. In ten weeks, the cost of inference per watt, the hardware tier on which models run locally and the reference model for half of all teams have all changed. An architecture that takes a specific supplier for granted is an architecture that has to be rewritten every cycle.
Frequently asked questions
What is Jalapeño and who makes it? It is the first inference accelerator designed by OpenAI, developed with Broadcom for the silicon implementation and the networking side. It was presented on 24 June 2026 and reached tape-out in nine months. Deployment in OpenAI's infrastructure is planned by the end of 2026, in small initial volumes.
What do the first Jalapeño benchmarks say? Measured on SemiAnalysis's InferenceX benchmark on GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5, the results published on 25 August 2026 indicate 1.5 to 1.9 times more peak AI work per kilowatt and 1.7 to 3.6 times lower end-to-end latency compared with the reference systems.
Is Jalapeño faster than Nvidia chips? On the published tests, yes, but the comparison is with the Blackwell generation currently in the field, not with Rubin. The data comes from OpenAI, the hardware is not accessible to third parties and nobody has independently re-run the tests.
Can I buy or rent Jalapeño? No. It is not a product sold on the market: it is internal silicon that OpenAI will use to serve its own models. The effect for developers will show, if at all, in the cost, latency and availability of the APIs.
Why is Claude Opus 5 criticised if it does well on benchmarks? Because the problem concerns the user experience, not correctness. The model produces roughly twice as many output tokens as its predecessors, with unrequested explanations, prefaces and checks. The verdict improves sharply when it works autonomously on verifiable goals and gets worse in conversational use.
Has Anthropic acknowledged the Opus 5 problem? The behaviour was documented in the prompting guide from the 24 July 2026 launch. After weeks of criticism came a public acknowledgement from a company engineer and, on 20 August, a "Concise" output style built into Claude Code. The structural fix is expected in later models.
How do you reduce Opus 5's verbosity today? By enabling the "Concise" output style and, above all, by removing inherited instructions like "double-check your work" from your prompts: on this generation they are redundant and amplify the problem. Lowering the effort reduces reasoning, not the length of the answer.
What changes with the Apple M6 chip? It is Apple's first 2-nanometre chip, with a 12-core CPU and GPU, a dual 16-core Neural Engine and up to 170 GB/s of unified memory bandwidth. In the new Mac mini Apple claims up to 4 times the AI performance of the previous generation.
How powerful is M5 Ultra? It is the most powerful chip Apple has ever made and its first quad-die: two dual-die M5 Max joined via UltraFusion, up to 36 CPU cores, 80 GPU cores, 1.2 TB/s of memory bandwidth and up to 512 GB of unified memory. It is exclusive to the new Mac Studio, available from 22 September 2026.
Will Apple launch M6 Pro, M6 Max and M6 Ultra? According to Bloomberg's reports, no: those variants would be cancelled, not postponed. The high end would move directly to the M7 family, with the base model in the first half of 2027, Pro and Max at the end of 2027 and Ultra in 2028. These are reports, not official Apple announcements.
Sources
- OpenAI: Jalapeño's first results show industry-leading speed and efficiency in AI inference
- OpenAI and Broadcom: OpenAI and Broadcom unveil LLM-optimized inference chip
- TechCrunch: OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show
- Tom's Hardware: Broadcom and OpenAI unveil custom-built Jalapeño inference processor
- Thariq Shihipar, Anthropic: statement on Opus 5 on X
- Anthropic: Prompting Claude Opus 5
- Anthropic: Prompting best practices
- Apple Newsroom: Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute
- Apple Newsroom: Apple unveils a more powerful Mac mini featuring the all-new M6 and M5 Pro
- Bloomberg: Apple to Skip High-End M6 Mac Chips, to Launch M7 Pro, M7 Max, M7 Ultra Instead


