Research

Towards rigorous, reliable, and capable AI scientists

  1. The Last Human-Written Paper: Agent-Native Research Artifacts

    What if a paper were something an AI could run and verify, not just read?

    arXiv · 2026

    arXiv · 2026ARA protocol

    The Last Human-Written Paper: Agent-Native Research Artifacts

    What if a paper were something an AI could run and verify, not just read?

  2. The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing

    Discovery is as sparse a signal as a crash, so what plays the role of coverage for an AI scientist?

    arXiv · 2026

    The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing

    Discovery is as sparse a signal as a crash, so what plays the role of coverage for an AI scientist?

  3. Sci-Reasoning: A Dataset Decoding AI Innovation Patterns

    How do the best researchers actually arrive at their breakthroughs, and could an AI learn to do the same?

    arXiv · 2026

    arXiv · 2026

    Sci-Reasoning: A Dataset Decoding AI Innovation Patterns

    How do the best researchers actually arrive at their breakthroughs, and could an AI learn to do the same?

  4. EXP-Bench: Can AI Conduct AI Research Experiments?

    Today's agents can design and code an experiment, so why can almost none of them finish one?

    ICLR · 2026

    ICLR · 2026

    EXP-Bench: Can AI Conduct AI Research Experiments?

    Today's agents can design and code an experiment, so why can almost none of them finish one?

  5. Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

    How do you make an AI scientist's experiments rigorous enough to trust the results?

    arXiv · 2025

    arXiv · 2025

    Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

    How do you make an AI scientist's experiments rigorous enough to trust the results?