What's new

What changed in our knowledge and in how we test ourselves.

  1. 10/10/2026 · monitor run
    Weekly monitor proposed 4 new sources
    Awaiting owner approval; nothing enters the corpus before that.
    Open source no link available
  2. 10/10/2026 · benchmark
    Benchmark release r4 (all split, 60 questions)
    Gates: 19/19 passed. Groundedness 94%, abstention 93%.
  3. 10/10/2026 · benchmark
    Benchmark release r3 (all split, 60 questions)
    Gates: 15/17 passed. Groundedness 92%, abstention 88%.
  4. 10/9/2026 · prompt
    Answer prompt v2 + multi-query retrieval (release r2)
    Added rules for tier-conflict weighting, false premises and off-topic abstention; retrieval now expands the question into up to 3 queries and caps 2 chunks per source; venue shown to the model; abstain threshold 0.33 → 0.36; relevance judge prompt clarified (judges-v2). See the r1 → r2 comparison on the Benchmark page.
  5. 10/9/2026 · model
    Models pinned: answers gpt-4.1-mini (OpenAI), judges gemini-2.5-flash (Google)
    Different providers to reduce self-preference bias. Judges are uncalibrated until the owner labels ~100 answers.
  6. 10/9/2026 · benchmark
    Benchmark release r2 (all split, 60 questions)
    Gates: 17/17 passed. Groundedness 96%, abstention 90%.
  7. 10/9/2026 · benchmark
    Benchmark release r1 (all split, 60 questions)
    Gates: 7/7 passed. Groundedness 98%, abstention 90%.
  8. 10/9/2026 · monitor run
    Weekly monitor proposed 8 new sources
    Awaiting owner approval; nothing enters the corpus before that.
    Open source no link available
  9. 10/9/2026 · corpus initialised
    Starter corpus loaded (33 sources)
    T1 papers, T2 official guidance, T3 practitioner writing; every source links to its original page.