Field Notes

Observations from building systems in the gap between demos and reality.
  1. Why confident models still make bad decisions

    Note
    001
    Status
    In preparation
  2. The cost of asking one more question

    Note
    002
    Status
    In preparation
  3. When should an AI refuse to continue?

    Note
    003
    Status
    In preparation
  4. What benchmarks don’t tell us

    Note
    004
    Status
    In preparation