Research
Four questions we keep returning to.
Reliability
Knowing when not to answer.
A system that always answers will sometimes answer wrongly with full confidence. We study when abstaining, asking, or deferring is the more useful output.
Context
Understanding what actually matters.
Real tasks arrive with partial, stale, or contradictory context. We look at how systems decide what to trust and what is missing.
Verification
Making systems check their own work.
Checking an answer is often easier than producing one. We explore how verification can sit inside a system rather than after it.
Failure
Learning from where models break.
Failures are data. We collect and categorise the ways systems break in real work, so they can be anticipated instead of discovered.
Reality and the model
- Observed points only
- Gap filled in
- 91%
- Wrong