M2-033M2: Classical Machine Learning梯度提升 (GBDT/XGBoost)Medium
Mastery:
梯度提升 (GBDT/XGBoost): XGBoost 相比传统 GBDT 有哪些改进?
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 二阶泰勒近似 + 显式正则项 + 加权分位点分裂 + 稀疏感知 + 列块并行。
📌 Key Takeaways
- •二阶信息使收敛更快更准
- •正则项 Ω 含叶子数与叶子权重 L2
📐 Mathematical Derivations
五项改进的具体内容:① <strong>二阶泰勒展开</strong>——目标函数展开到二阶,用梯度 gᵢ 与海森 hᵢ,得到叶子权重的<strong>闭式最优解</strong> w*_j=−G_j/(H_j+λ) 与相应的最优增益公式。二阶信息使每步更新更准(类似牛顿法 vs 梯度下降),收敛更快;同时海森 hᵢ 自然地充当<strong>样本权重</strong>(hᵢ 大的样本影响更大),这解释了 XGBoost 对不平衡数据的处理能力。② <strong>显式正则项</strong> Ω(f)=γT+½λ‖w‖²——惩罚叶子数 T(控制复杂度)与叶子权重(防过拟合),把正则内建到目标中而非事后剪枝。③ <strong>加权分位点分裂(weighted quantile sketch)</strong>——对连续特征用 hᵢ 加权的分位点找候选分裂点,兼顾精度与效率(O(#bins) 而非 O(#unique))。④ <strong>稀疏感知</strong>——为缺失值/稀疏值学习默认方向。⑤ <strong>列块并行 + 缓存优化</strong>——特征按列压缩存储,分裂搜索可多线程并行(注意:树的生长仍是串行的,并行在特征维度)。
🏭 Production Trade-offs
实践要点:① <strong>核心超参</strong>——<code>eta</code>(学习率)、<code>max_depth</code>(默认 6)、<code>min_child_weight</code>(叶子最小海森和,控制过拟合)、<code>subsample</code>(行采样)、<code>colsample_bytree</code>(列采样)、<code>lambda/alpha</code>(L2/L1 正则)、<code>gamma</code>(分裂最小增益)。② <strong>gamma 的作用</strong>——分裂增益必须超过 γ 才分裂,是一种<strong>预剪枝</strong>;γ 越大模型越保守。③ <strong>与 LightGBM 的差异</strong>——LightGBM 用直方图算法(把连续特征分桶,O(#bins))与 <strong>leaf-wise 生长</strong>(每次分裂增益最大的叶子,收敛快但需限制深度/叶数防过拟合),并加 GOSS(梯度单边采样)与 EFB(互斥特征绑定)进一步加速。④ <strong>实践建议</strong>——表格数据上 XGBoost/LightGBM/CatBoost 通常最强;优先调 <code>eta</code> 与树数、再调深度与正则;用早停(<code>early_stopping_rounds</code>)防止过拟合。
⚠️ Common Interview Pitfalls
- ✕只用一阶梯度而不利用二阶信息
- ✕不设 min_child_weight 导致叶节点样本过少而过拟合
🎯 Interviewer Follow-ups
- ?为什么二阶泰勒优于一阶?
- ?XGBoost 如何做并行?(特征维度并行)
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.