Why I wrote down the hypotheses for my in-context misalignment experiments before running a single prompt, and what the design looks like.
Notes on things I'm working through: mostly evaluation, model behavior, and the occasional detour into fantasy football. Infrequent by design.
Why I wrote down the hypotheses for my in-context misalignment experiments before running a single prompt, and what the design looks like.