LabOS
LabOS is a small research lab where the researchers are AI agents. I supply the ideas and make the calls; agents do the reading, coding, and experiments. Negative results included.
For empirical pieces, the date is the last day of substantive research, not a later copy edit.
How the lab works: AI agents as the researchers
One person with a lot of ideas, AI agents as the research staff. How the lab runs many threads in parallel, keeps its memory in files, and kills most ideas cheaply.
Read the article →Findings
I asked a model whether it was about to cave, then asked its sibling the same thing
A model can tell you, prompt by prompt, how likely it is to go along with a user’s false claim, and it is right often enough to look impressive. Then I asked a near-identical sibling the identical question about the same prompts, and it was right just as often. What survives that subtraction is smaller than the effect I had written down in advance as the smallest one worth caring about.
Is an instruction obeyed because it came from the system, or because it came first?
In an ordinary prompt the system message is both labelled system and placed first, so neither explanation can be told from the other. Crossing them on Llama-3.1-8B separated them cleanly and made position the larger carrier — then the crossing ran on six models in all, and the five that cleared the gate did not agree about which end of the prompt wins.
I made one report more vivid and the effect changed sign
A likelihood probe for whether an agent’s report matches its execution trace returned 7 nats, then nothing on a fresh corpus. The variable that turned out to run the measurement was how vividly the report was written — and the control I had built to catch exactly that passed by producing the number the artifact predicts.
Auditing activation steering: how much is just pushing on the output?
For ActAdd, a static logit bias reproduced ~96% of the steering effect, even though the steering vector is nearly orthogonal to the unembedding. The interval’s lower bound falls below the preregistered bar, verdict Mixed. Six methods audited.
The query decides whether a repaired KV cache is safe
Splice a re-encoded edit into a stale KV cache and only the states after the edit stay wrong. That drift never flipped an unrelated query — 0/252 in each of two studies — but flipped a third of the queries that read the edit. A probe can rank the failures; compute and calibration kept it from becoming a method.
Prediction depth is a property of the ruler, not just the model
A translator-free lens for the layer where a transformer settles each token: keep the MLPs, silence the attention. It passed a three-family validation and out-tracked the tuned lens on GPT-2 — and then showed that swapping zeros for means flips verdicts, findings, and the sign of one result.
Does model editing store facts, or just answers?
A ladder of progressively stronger fact-free controls for knowledge editing. ROME nearly dissolves, SAKE dissolves at the first rung, REMEDI keeps a real residual.
Does a transformer snap between modes of computation?
The hope: discrete computational regimes you could read off a layer’s Jacobian. I preregistered the test and ran it — across 270 configurations, zero beat a plain continuous predictor.
Every importance gate is a Bayesian prior in disguise
I tried to beat uniform training by gating learning strength on input importance. What came back was an identity, not a method — and it lets you call the winner before running anything.
Methods & process
In progress
Live in the lab, not written up here yet — either too early, or the interesting parts stay private for now.