Back to News
RSS feedwww.lesswrong.com

Interactive Map Visualizes the AI Safety Research Field

Summary

A LessWrong post presents an interactive terrain-style visualization of about 3,466 AI safety works collected from arXiv and LessWrong-related sources through a keyword dictionary. The map clusters the works in embedding space across 18 subfields, with terrain height based on log-compressed citation counts aggregated across neighboring works. Users can explore the field over time from 2021 Q1 to 2026 Q3, filter by citation count or researcher, and search individual works; the dataset includes 3,989 named authors. The terrain contains papers only because Semantic Scholar does not index LessWrong, while forum material is used for researcher profiles and is planned for a separate filtered dataset. The visualization suggests that alignment training and scalable oversight currently have the highest citation concentration, followed by adversarial robustness and interpretability. Alignment and scalable-oversight citations are concentrated in a small number of highly cited 2022-23 papers, including InstructGPT, DPO, Anthropic's helpful-and-harmless RLHF work, and Constitutional AI. By contrast, adversarial robustness and interpretability contain more works but lower citations per paper. The author cautions that citations are a skewed and imperfect proxy for impact: the top 1% of works account for 40% of citations in the set, and citation feedback loops may not reflect actual safety contributions. The project remains a work in progress and used Claude extensively for scraping and embedding-analysis code.