AI Roadmap/Layer 07 · 07. AI Platform / MLOps / LLMOps
7.3

7.3 Evaluation, PromptOps & Quality Gates

Prompt-as-code versioning, environment config injection, golden benchmark datasets, automated LLM-as-a-judge calibration, and regression gates.

Prompt Registry & PromptOps Lifecycle Management

Prompt-as-code management, dynamic parameter templating, version tagging, GitOps review workflows, A/B prompt experiment testing, and team collaboration workspaces.

🏢 Companies
Weights & BiasesBraintrustLangfuseHumanloopPromptfoo
🛠️ Tech Stack
PromptOpsPrompt RegistryPrompt-as-CodeA/B TestingGitOpsBraintrustHumanloop
💼 Roles & Salary
LLMOps Engineer、QA Automation Engineer、AI Platform Engineer
💰 $175K - $360K / year (PromptOps & Prompt Management) | ¥450K - ¥1.05M / year
📚 Prerequisites: Prompt Templating Engines (Jinja2) • GitOps & CI/CD Pipeline Automation • Prompt Injection Defense & Sanitization • A/B Testing & Statistical Analysis

Golden Datasets, LLM-as-a-Judge & Automated Eval CI/CD

Golden evaluation dataset curation, multi-criteria LLM-as-a-judge scoring calibrated against human review, statistical significance testing (Pass@k), and automated pre-deployment regression gates.

🏢 Companies
BraintrustLangfuseArize AIWeights & BiasesPromptfoo
🛠️ Tech Stack
LLM-as-a-JudgeEval CI/CDGolden DatasetRegression TestingStatistical SignificanceBraintrustArize AI
💼 Roles & Salary
Evaluation Engineer、QA Automation Engineer、Applied ML Engineer
💰 $185K - $380K / year (Automated AI Evaluation & CI/CD) | ¥480K - ¥1.15M / year
📚 Prerequisites: Eval Frameworks (DeepEval/Promptfoo) • LLM-as-a-Judge Bias Mitigation • Golden Test Suite Curation & Coverage • Statistical Significance Testing (Bootstrap/t-test)