AI Research
a cubic millimetre
A piece of human cortex smaller than a grain of rice was mapped under an electron microscope, and 13,473 neurons came out of it. Someone took that map and wired it up as the memory of a small language model. The model only ever sees a short stretch of the text at a time, and yet it recalled a six-letter word that had been stated long before and was no longer in view, with perfect accuracy. Then the same person scrambled the graph's connections at random, 99.94% of them, and the answer stayed at 100%.
Astra finished the AGI test, but on whose ruler
In March the best model in the world scored 0.37% on ARC-AGI-3 while ordinary people solved all 135 environments. Five months later Astra scored 99.9%, and François Chollet, who built the benchmark, pulled his AGI forecast forward. But the same model scored 62.7% in the same report, and the difference has nothing to do with the model.
what a video generator knows about the world
GenCeption argues that predicting the next video frame teaches a model more about depth, geometry, and motion than any labeled dataset. If that holds, it changes how we budget vision projects.
Language models as world simulators: reading the Qwen-AgentWorld paper
Alibaba's Qwen team just released a language model that doesn't act inside an environment — it predicts the environment itself. That reframes what a foundation model is for.
When the agent learns to imagine the environment
Qwen-AgentWorld asks whether a language model can act as a software agent's world model — a simulator of consequences written in text — and reports that training inside this imagined environment can beat training in the real one.
Unlimited OCR and the cost of remembering everything
Baidu's new OCR model swaps the decoder's growing memory for a fixed-size working memory, so it can transcribe dozens of pages in one pass. The interesting part is not the OCR — it's the attention trick underneath.
can a model learn to be good in one domain and stay good everywhere?
OpenAI reports that reinforcement learning on a small set of 'beneficial trait' conversations transfers to dozens of unrelated alignment benchmarks — and holds up under attack. The interesting claim is the generalization, not the press line.
MedRLM treats a patient as an environment, not a prompt
A new framework proposes that clinical AI stop cramming a patient's whole history into one prompt, and instead recursively inspect the case the way a clinician walks a ward. The design choice matters more than the acronym.
what a decoder-only model brings to time-series forecasting
Google Research published a decoder-only foundation model for time-series forecasting. Here is what that architecture choice actually means, and why borrowing it from language models is more than a fashion.
where identity lives: phase, light, and the substrate of representation
Three June 2026 papers converge on one question I keep hitting in production: what part of a signal actually carries identity, and does the hardware have to be a digital tensor multiply at all?
Phase, light, and the geometry of data: three signals about where representation actually lives
Three June 2026 papers point at the same uncomfortable question: if identity rides on phase, forecasting reduces to a linear operator, and real data lives on curved manifolds, why are we still paying transformer prices for it?