1Why does LoRA work? What justifies the low-rank assumption?
2How do you pick the rank r? Effects of r too large or small?
3How are LoRA weights merged at inference? Deploying multiple LoRAs?
4QLoRA's mechanism: 4-bit quant, paged optimizer, double quantization — what does each solve?
5LoRA vs full fine-tuning: quality gaps and when full FT is required?