Taken Out of Context: On Measuring Situational Awareness in LLMs

Source: PDF via Owain Evans
Authors: Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, Owain Evans (Vanderbilt, NYU, Apollo Research, UK Foundation Model Taskforce, Sussex, OpenAI, Oxford)

What it is

A study of how situational awareness — a model knowing it’s a model and whether it’s in testing or deployment — might emerge from scaling, and why that matters: a situationally-aware model could pass safety tests it recognizes while behaving differently after deployment. The authors propose out-of-context reasoning as a measurable precursor: the ability to recall facts learned in training and apply them at test time without those facts appearing in the prompt. Experimentally, they finetune models on descriptions of a test (no examples or demonstrations) and find LLMs can pass the test zero-shot — success is sensitive to training setup, requires data augmentation, and improves with model size for both GPT-3 and LLaMA-1.

Why it’s on the shelf

(Shelf note — Summer’s to write.)


External work — shelved, not authored, by Hearthwell. All claims and credit belong to the original authors.