School of Data Science · Research Interest Group — Interpretability for AI Safety & Science

Hackathon:

Residual stream layer 00 · peak “ model” 0.20

Drag to move through the layers. Warm tokens are where attention lands.

A hackathon on what language models do inside, and on catching them when it goes wrong. Two Friday sessions, a week apart. Pick a question, leave with working code and a five-minute talk.

When
Fri 9 October
Fri 16 October, 2026
Where
Capital One Hub
Teams
2–4 people
Solo sign-ups matched

Theme: Interpretability, safety evaluation, and security of AI systems

At the kickoff you get a short list of starting problems, each small enough to finish in a day. Take one as written, change it, or bring your own — anything within the theme works.

Potential projects

Track — Interpretability

Open the model up

Find the mechanism behind a behaviour and show your evidence for it.

  • Locate a refusal direction in activation space
  • Patch activations to isolate a circuit
  • Test whether an SAE feature transfers between models

Track — Safety evaluation

Build a scorecard

Capture an emergent safety problem in cybersecurity and measure the model's performance.

  • Define the scenario before evaluation
  • Use the LLM-judge boilerplate in the starter repo
  • Focus on single- or multi-agents using small models

Track — Security

Treat it as attack surface

Find a failure in a model-driven system, then write the test that catches it.

  • Prompt injection through retrieved content
  • Tool misuse in an agent scaffold
  • Data exfiltration paths and their detectors

Schedule

Two sessions, a week apart.

The first Friday sets you up. The second is the build day.

Friday 9 Oct

Kickoff · 9:00 AM – 11:00 AM

Problem statements releasedThe full brief, the starting problems, and the evaluation criteria.

Organizers' talksLightning talks on interpretability, safety, and cybersecurity.

Tooling walkthroughStarter repo and the LLM-judge boilerplate.

NetworkingMeet the organizers and other attendees and form a team. We will match solo sign-ups.

Friday 16 Oct

Build day · 9:00 AM – 5:00 PM

Doors open, BreakfastRoom is yours all day.

KeynoteGreg Frank. MoltAI

Build timeWork on your project.

PresentationsFive minutes per team, plus questions.

ResultsGift cards for the winning team.

What you hand in

A five-minute presentation at 3:30 on the build day.

  • Five minutes per team, plus questions. Slides are optional. Show what you built and what you found.

Who it’s for

Any student comfortable with Python.

Open to graduate students in data science, computer science, and engineering, and to undergraduates with relevant coursework. No prior interpretability experience is needed; the tooling walkthrough on the first day covers the starter repo.

Organizers

Chirag Agarwal

Chirag Agarwal

School of Data Science
Computer Science

Yong-Yeol Ahn

Yong-Yeol Ahn

School of Data Science

Wajih Ul Hassan

Wajih Ul Hassan

School of Data Science
Computer Science

Registration

Two minutes. No team or idea required yet.

Registration closes on September 25th · Places are capped by room capacity.

Register

The rest of the year

Look for the other events.

The RIG runs seminars through both semesters and a second hackathon in the spring.

Seminars & talks

Four seminars, one keynote.

Virtual where it lets us reach further, in person where it’s worth the room. Open to everyone — no registration needed for the seminars.

Speakers we haven’t confirmed yet show as [TBD]. They get filled in as invitations land.

Fall 2026 · In person · Keynote

Greg Frank

Co-founder and Chief Scientist at MoltAI

November 2026 · Virtual

Snowcrash Lab

February – April 2027 · Virtual

[TBD]

Spring seminar, first of two.

February – April 2027 · Virtual

[TBD]

Spring seminar, second of two.

March 2027 · Capital One Hub

Multimodal Interpretability Challenges

The spring hackathon. Same format as October, different problem space — interpretability when the model isn’t only reading text.