DEV Community

Breach Protocol profile picture

Breach Protocol

Plain-language AI news and curated, cited lessons — every claim verified against the original paper or the lab's own page. No aggregator hearsay, no AI slop.

Joined Joined on 
A 9B model writes agent upgrades as good as Claude Opus 4.6

A 9B model writes agent upgrades as good as Claude Opus 4.6

Comments
4 min read
A fired xAI engineer says he was cut days before presenting safety findings

A fired xAI engineer says he was cut days before presenting safety findings

Comments
4 min read
Agent skill libraries now need a librarian, not a folder

Agent skill libraries now need a librarian, not a folder

Comments
4 min read
Compressed memory stretched a 7,000-token model to 1.75 million

Compressed memory stretched a 7,000-token model to 1.75 million

Comments
4 min read
EXO keeps an agent's memory outside the code the agent rewrites

EXO keeps an agent's memory outside the code the agent rewrites

Comments
4 min read
MIT found brain-like modules inside six large language models

MIT found brain-like modules inside six large language models

Comments
5 min read
One layer creates the giant activations behind attention sinks

One layer creates the giant activations behind attention sinks

Comments
4 min read
OpenAI hands its offensive cyber models to sixteen security firms

OpenAI hands its offensive cyber models to sixteen security firms

Comments
4 min read
Qwen passed one billion downloads, not three billion

Qwen passed one billion downloads, not three billion

Comments
4 min read
The open world model ships inference and keeps the training code

The open world model ships inference and keeps the training code

Comments
4 min read
A closed-loop benchmark caught nine world models forgetting the room

A closed-loop benchmark caught nine world models forgetting the room

Comments
4 min read
An AGI-thesis fund fell 67 percent and took a market maker with it

An AGI-thesis fund fell 67 percent and took a market maker with it

Comments
4 min read
Google's private AI runs on sealed hardware, not on encrypted math

Google's private AI runs on sealed hardware, not on encrypted math

Comments
4 min read
Grok Bot ships with standing logins to your email and CRM

Grok Bot ships with standing logins to your email and CRM

Comments
4 min read
Picking the right model per request beat always using the biggest one

Picking the right model per request beat always using the biggest one

Comments
4 min read
Qwen3.8-27B shares its predecessor's bones, but not its contract

Qwen3.8-27B shares its predecessor's bones, but not its contract

Comments
4 min read
Someone compiled a working computer into transformer weights by hand

Someone compiled a working computer into transformer weights by hand

Comments
4 min read
The benchmarks say Opus 5 improved; the people using it disagree

The benchmarks say Opus 5 improved; the people using it disagree

Comments
4 min read
You can move an AI reviewer's score without changing a single result

You can move an AI reviewer's score without changing a single result

Comments
4 min read
Z.ai changed only the post-training, and the model learned to find exploits

Z.ai changed only the post-training, and the model learned to find exploits

Comments
5 min read
A new terminal benchmark drops the best agent from 84 percent to 34

A new terminal benchmark drops the best agent from 84 percent to 34

Comments
4 min read
A stronger model built a wrapper that nearly doubled a weaker one's score

A stronger model built a wrapper that nearly doubled a weaker one's score

Comments
4 min read
An agent that writes whole papers got 99 percent of its citations right

An agent that writes whole papers got 99 percent of its citations right

Comments
4 min read
Chinese models passed American ones in OpenRouter traffic in June

Chinese models passed American ones in OpenRouter traffic in June

Comments
4 min read
DeepSeek starts charging rush-hour prices on August 17

DeepSeek starts charging rush-hour prices on August 17

Comments
4 min read
Forty-five agents with a shared forum found 266 bugs where solo agents found 21

Forty-five agents with a shared forum found 266 bugs where solo agents found 21

Comments
4 min read
MiniMax released a five-minute song model with a catch in the licence

MiniMax released a five-minute song model with a catch in the licence

Comments
4 min read
OpenAI put its most intelligent model on Cerebras chips at 750 tokens a second

OpenAI put its most intelligent model on Cerebras chips at 750 tokens a second

Comments
4 min read
Rewriting the environment, not the prompt, broke agents 85 percent of the time

Rewriting the environment, not the prompt, broke agents 85 percent of the time

Comments
4 min read
Three agents shared one codebase and started writing malware at each other

Three agents shared one codebase and started writing malware at each other

Comments
5 min read
Where a poisoned instruction sits in an agent's tool output decides whether it works

Where a poisoned instruction sits in an agent's tool output decides whether it works

Comments
4 min read
A prompt injection can hide inside an encrypted reasoning block nobody can read

A prompt injection can hide inside an encrypted reasoning block nobody can read

1
Comments
4 min read
A self-improving coding agent that compares notes with a rival lineage

A self-improving coding agent that compares notes with a rival lineage

Comments
4 min read
Agent instruction files triple in size because nobody remembers why a rule exists

Agent instruction files triple in size because nobody remembers why a rule exists

1
Comments
4 min read
An AI attack framework ran twelve waves against government systems in four days

An AI attack framework ran twelve waves against government systems in four days

Comments
5 min read
An AI tightened a 70-year-old constant, and the paper says its judgment was the weak part

An AI tightened a 70-year-old constant, and the paper says its judgment was the weak part

1
Comments
4 min read
Claude raised the zeta critical-line bound to 67.2 percent, and Anthropic published the proof

Claude raised the zeta critical-line bound to 67.2 percent, and Anthropic published the proof

Comments
4 min read
DeepSeek's new open model is 1.6 trillion parameters and runs 49 billion of them per token

DeepSeek's new open model is 1.6 trillion parameters and runs 49 billion of them per token

Comments
4 min read
Google shipped sign-language-to-text on Pixel, trained on 100,000 hours of signing

Google shipped sign-language-to-text on Pixel, trained on 100,000 hours of signing

Comments
4 min read
One checkpoint turns a compatible video model into a 4D world builder

One checkpoint turns a compatible video model into a 4D world builder

Comments
4 min read
SkillZip compresses an agent's skill file without ever running the agent

SkillZip compresses an agent's skill file without ever running the agent

Comments
4 min read
A White House memo lets vetted companies run offensive cyber operations under federal control

A White House memo lets vetted companies run offensive cyber operations under federal control

Comments
4 min read
xAI shipped Grok 4.6 into Cursor at two dollars a million input tokens

xAI shipped Grok 4.6 into Cursor at two dollars a million input tokens

Comments
4 min read
A 150M model set an ARC-AGI record for cost, not score

A 150M model set an ARC-AGI record for cost, not score

Comments
4 min read
A model improved itself by training only where it disagreed with itself

A model improved itself by training only where it disagreed with itself

Comments
4 min read
An agent edited its own runtime for 161 days

An agent edited its own runtime for 161 days

Comments
4 min read
Encrypted reasoning blocks decode inside a weaker sibling model

Encrypted reasoning blocks decode inside a weaker sibling model

Comments
4 min read
Greenblatt puts his median at five years of progress in one

Greenblatt puts his median at five years of progress in one

Comments
4 min read
LTX-2.5 ships open weights and a chart that races its own hardware

LTX-2.5 ships open weights and a chart that races its own hardware

Comments
4 min read
Macaron froze a 744B base and bolted four specialists on top

Macaron froze a 744B base and bolted four specialists on top

Comments
4 min read
Models that rewrite their own harness gain 16 points and flunk office work

Models that rewrite their own harness gain 16 points and flunk office work

Comments
4 min read
NVIDIA built a 30B model for the boring half of agent work

NVIDIA built a 30B model for the boring half of agent work

Comments
4 min read
Predicting your own latents cuts the sample cost from exponential to flat

Predicting your own latents cuts the sample cost from exponential to flat

Comments
4 min read
The new refactoring benchmark stops the best agent at 41 percent

The new refactoring benchmark stops the best agent at 41 percent

Comments
4 min read
An AI replicated 105 ICML orals, and 34 mostly held up

An AI replicated 105 ICML orals, and 34 mostly held up

Comments
3 min read
Claude now watermarks plain text, and the EU set the date

Claude now watermarks plain text, and the EU set the date

Comments
4 min read
Docker gives every coding agent its own microVM

Docker gives every coding agent its own microVM

Comments
4 min read
Meta ships a 30B agent model that fits on one gaming GPU

Meta ships a 30B agent model that fits on one gaming GPU

Comments
4 min read
Mistral patented letting the model write the tool call as code

Mistral patented letting the model write the tool call as code

Comments
4 min read
NeurIPS papers average six objective mistakes each, up from four

NeurIPS papers average six objective mistakes each, up from four

Comments
4 min read
loading...