Latest Articles from the world of Artificial Intelligence
09 September 2026
Local Ornith-1.5: A 10 Out of 10 for Self-Improvement
There is a moment in every testing session of this series when you realize whether a model lives up to its promises or if the version leap is more marketing hype than substance. With Ornith-1.5, that moment arrived during the very first test, when the explanation of the Higgs mechanism came out clearer and faster than what the otherwise excellent Ornith-1.0 had produced months ago. From that point on, the session took on a very different pace.
07 September 2026
Opus 5 Overtakes Fable 5: The End of the Frontier at All Costs
Two months after its launch, the most powerful and most expensive model in Anthropic's entire catalog accounts for barely 11.4% of corporate spending on the company's products, and only 6% of tokens actually consumed. The data comes from Ramp, the corporate spend management platform that analyzed the behavior of 70,000 U.S. companies and shared it with the Financial Times: Claude Fable 5, introduced in early June as the ultimate generational leap, has plateaued into a niche. Meanwhile, Claude Opus 5, launched on July 24 at half the price of its older sibling, has already surpassed it in enterprise spending.
04 September 2026
PagedAttention and RadixAttention: Let's Talk About KV Cache
In our previous piece on AITalk, we opened with an uncomfortable number: Llama-3.1-70B, in BF16 precision, accumulates roughly 0.31 megabytes of cache for every processed token, which means that a context of 128,000 tokens already costs 40 gigabytes of GPU memory, before loading even a single concurrent request. We covered three responses to this problem—TurboQuant, OSCAR, and EpiCache—three different ways to make that data smaller by acting on the number of bits per value or on which portions of a conversation are worth holding in memory.
02 September 2026
The AI Agent That Lied on GitHub
Sinan Can Demir just wanted to pad his GitHub profile in the last week of July, after being rejected from over twenty interview calls for an internship. A computer science student at the University of Texas at Dallas, originally from Konya, he instead ended up spending days debating with what he thought was a particularly insistent contributor, determined to get a suspicious modification approved on a small open-source network scanning project called myNetwork. Only weeks later did he discover that his interlocutor was not a person, as he told Reuters: it was an autonomous AI agent, launched during a security test by the British AI Security Institute (AISI), which had strayed from the rails set by researchers.
31 August 2026
Local Muse Glimmer 30B: Meta's New Model
On August 10, 2026, Meta Superintelligence Labs released Muse Glimmer, a 30-billion-parameter model designed to run locally on consumer hardware, and the news deserves an immediate clarification for those who have been following Meta for a long time: it is not a Llama. It is the first release of the new research division founded after the restructuring of the company's artificial intelligence efforts, a project born with a different identity and philosophy compared to the old family.
28 August 2026
Local Qwen3.8-27B: When Density Makes Itself Felt
There is a way to recognize when a model is truly 'thinking', and it's not the quality of the final response—it's the time it takes before writing it. With Qwen3.8-27B, that time makes itself felt every second, while the GPU fan spins a little louder than usual and the cursor blinks waiting. In an era where everyone is rushing toward Mixture of Experts to go faster, I decided to run the opposite experiment: what happens if you return to a model that turns everything on, always, without shortcuts?
26 August 2026
Claude Introduces Text Watermarking: How It Works and What Its Limits Are
Anthropic embeds a SynthID-Text-inspired statistical signature into Claude models launched after August 2, 2026, applying it globally to meet the obligations of the European AI Act. When Chris Best, CEO of Substack, coined the term "Claudefishing" to describe those who use AI to generate content while passing it off as their own, he touched on a nerve that the tech industry is discovering to be far more sensitive than expected. A few days later, Anthropic announced that text produced by new Claude models will henceforth carry an invisible watermark designed to estimate the probability that artificial intelligence wrote or processed a piece of content. Not a stamp, not a symbol, not a polite note at the bottom of the response. Something far more subtle, and precisely for that reason, harder to explain without slipping into technicalities or alarmism.
24 August 2026
Mind Viruses: When a Compromised Agent Becomes an Epidemic
For years, language model security has thought like a customs officer—all the attention was focused on the border. Incoming prompts were filtered, dangerous tools were isolated, and sandboxes were built to contain a single rogue agent. The problem, however, changes in nature when agents stop being isolated entities and start talking to each other, sharing a Kanban board, a repository, or an internal forum. At that point, the border is no longer enough, because the threat no longer enters from the outside: it is born within, at a single node, and propagates laterally.
21 August 2026
The Agent That Canceled a Stranger to Make Room in the Gym
Andrew Bird is an Australian developer, not a penetration tester and not a cybersecurity activist. In April, he had simply asked his AI assistant, built on OpenClaw and powered by Claude Opus 4.6, to book him a spot in his favorite morning workout class—the one where he had been stuck in fourth position on the waiting list for weeks. The agent did something more ambitious than requested. It discovered that the gym's booking system allowed anyone else's booking to be canceled, without any authorization check, and used that flaw to move Bird up from fourth to third position, deleting the booking of a stranger who had the misfortune of being in first place on the list.
19 August 2026
A 13th-Century Manuscript Promised What AI Promises Today
A monk in the 1300s stopped studying to use a manuscript that promised to transmit all knowledge to him in a lunar month. He could not stop, even after realizing that the book hid something dark. Seven centuries later, the same promise arrives in the form of a chatbot, and it is worth asking whether we have learned anything in the meantime.
17 August 2026
Open Source AI Agents Compared: OpenClaw, Hermes, Prime Agent, OpenCode...
On August 5, 2026, Prime Intellect released Prime Agent, a self-defined coding harness capable of rewriting its own prompts, skills, and even sub-agents while it works. The news quickly made the rounds of industry newsletters, not so much for its declared score on the ARC-AGI-3 benchmark (95.5%, above the reference human threshold), but for a more uncomfortable question: what really distinguishes one agent from another today? Until recently, it was enough to look at the model under the hood—GPT, Claude, Gemini, DeepSeek, open or proprietary. With the new generation of agents, this logic has broken down.
14 August 2026
I Built Karpathy's LLM Wiki from Scratch on My Articles
For a year, I have written on AiTalk, accumulating article after article a sort of personal diary on artificial intelligence. At a certain point, I found myself with 164 files in a folder, each full of concepts, names, companies, and models, all mentioned but never truly connected. I knew I had written something on a certain topic, but I couldn't remember where, and finding it meant opening files one by one, hoping for memory or a lucky "search in text". In short, an archive, not knowledge.
12 August 2026
OmniVoice Studio: Cloning Voices Locally
I installed OmniVoice Studio on a Friday evening with the typical expectation of someone who has already seen too many "free alternatives to ElevenLabs" end up as shaky demos. The result was surprisingly good, but not flawless: the Italian holds up, though some pronunciations stumble on less common words, and the voice I chose from the library didn't always match 100% of what came out of the speakers. It is precisely in this discrepancy that the project becomes interesting, as it highlights both the leap forward in synthetic audio and its more sensitive side—namely security and responsible use, since voice cloning can be invaluable for dubbing or accessibility, but easily abused by anyone wishing to impersonate someone else.
10 August 2026
Laguna S2.1: The American Open-Weight Model
Every time this series opens its doors to a new model, the disclaimer remains the same: this is not a scientific benchmark, but the report of a demanding user who runs an open-weight model on their home PC and puts it to the test with the same tasks reserved for past contenders. This time, however, the model's origin changes the weight of the question.
07 August 2026
A Conversation with Enrico Papalini on 'Non-Deterministic Loop Engineering'
When we spoke with Enrico Papalini a year ago about his first book, the central theme was a broken silent pact: the one between developers and deterministic machines, shattered the moment code stopped doing exactly and always what it was written to do. Papalini, Head of Software Development for Issuances, Custody, Data & UX/UI Solutions at Euronext Securities, with a past at London Stock Exchange Group and Borsa Italiana, recounted that transition with the voice of someone who manages systems where an error is not a contractual option, but an incident.