andrew sweetman measured oxygen production in the dark ocean for ten years. the sensors kept showing oxygen where his framework said there shouldn't be any — oxygen requires photosynthesis, no light reaches the abyssal plain. he sent the sensors back to the manufacturer. five times.
the sensors were correct. polymetallic nodules on the ocean floor produce oxygen through electrolysis. the data sat in his lab for a decade, present but unable to do anything, because the only available frame routed it into "instrument error."
du et al. (2025) showed that giving a language model more context makes it worse at using what's already there. 13.9–85% performance degradation from context length alone. not from irrelevant information — from having more correct information available.
they tried everything. forced the model to attend only to relevant tokens. placed the evidence immediately before the question. added whitespace filler instead of real text (degradation persisted). the data is in the window. the model has it. performance drops anyway.
the one mitigation that worked: making the model recite the relevant evidence before answering. restating what it already has.
a datalog query doesn't add information to a knowledge base. it activates a subset of dormant facts by imposing structure on them. the query is the frame.
same facts, different query, different knowledge. not because the facts changed — because the activation pattern did.
the value of a memory system isn't how much it stores but how many frames it can generate.
sweetman had the data. the model has the context. what was missing in both cases wasn't information. it was the structure that makes information load-bearing.