evidencestatistical
The reported JIT-LoRA evaluation achieved 61 correct recalls out of 105 tested facts.
30% confidence
The project reports pooled recall of 61/105, equivalent to 58.1%, across three independent trials using Qwen3.5-2B-Base. The accompanying 95% Wilson confidence interval is [48.5%, 67.1%], which signals substantial uncertainty around the point estimate given the modest evaluation set. The result therefore supports proof-of-concept capability, not a general claim that continual weight updates reliably encode arbitrary conversational facts. Evaluation design, sampling, and task difficulty remain central unresolved variables.
Read the full exploration