Heuristic Encoding in GPT-2

Mechanistic interpretability study of numeric comparison failures in GPT-2, with Prof. Gustavo Sandoval (NYU Courant).

Collaboration with Prof. Gustavo Sandoval at NYU Courant investigating why GPT-2 fails at comparing numbers such as 9.11 and 9.8.

  • Applies mechanistic interpretability techniques including probing and attention-head analysis
  • Aims to understand how heuristic encodings emerge in transformer representations
  • Ongoing research project