SARSA vs. Q-learning on the cliff revisited

I treat the textbook cliff-walking comparison of SARSA and Q-learning as a controlled experiment — randomness isolated into named streams, the two algorithms paired with common random numbers — and separate behavior return from greedy-evaluation return.

Emergence as metric composition in LLMs (Part I): the parameter axis

Part I of two. Borrowing the finite-size-scaling methodology from statistical physics, I ask whether emergent abilities postulated to appear in larger models are a measurement artifact across eight Pythia sizes. Specifically, can a benchmark metric's sharpness be explained entirely from composing a smooth per-token accuracy through a hard metric or are there other signatures of emergence hidden in the per-token accuracy?