evidenceacademic
For standard transformer architectures, the effect of in-context learning cannot be converted exactly into compressed context stored in the model's parameters.
30% confidence
The cited analysis draws a sharp architectural boundary. It reports that exact conversion of context into model weights is mathematically impossible for standard architectures, while modified attention mechanisms with additional bias terms can support context compression in restricted settings. The same work identifies a narrower result: context effects can be represented as low-rank updates to feedforward-network weights. That is a functional equivalence, not evidence that ordinary inference permanently edits deployed parameters.
Read the full exploration