1Few-shot vs zero-shot CoT? Why does "Let's think step by step" work?
2How does CoT emergence relate to model size, and why do small models benefit little?
3How does self-consistency work, and why does majority voting improve accuracy?
4Limitations of CoT? When is it ineffective or even harmful?
5How does CoT relate to test-time compute scaling?