Back to AI Roadmap
📊
LAYER 03

03. Data Engineering & Data Infrastructure

Benchmark Companies:
Scale AIScale AILabelboxLabelboxSurge AISurge AISnorkel AISnorkel AIGretel.aiGretel.aiAppenAppenDatabricksDatabricksSnowflakeSnowflakeConfluentConfluentFivetranFivetrandbt Labsdbt LabsCollibraCollibraMonte CarloMonte CarloBigIDBigIDPineconePineconeQdrantQdrantZilliz / MilvusZilliz / MilvusWeaviateWeaviateClickHouseClickHouseElasticsearchElasticsearchOpenSearchOpenSearchVespaVespaHugging FaceHugging Face

Curating web-scale datasets, synthetic generation engines, enterprise data governance, and high-performance retrieval indexing.

📊Layer View Mode:
🌐3D CROSS-CORRELATION MATRIX

03. Data Engineering & Data Infrastructure · Cross-Correlation Ecosystem

Click any company, tech category, or career role to illuminate all cross-associations and dim unrelated entities.

🔍
🏢

Benchmark Enterprises

23
Scale AIScale AI
2
LabelboxLabelbox
1
Surge AISurge AI
2
Snorkel AISnorkel AI
1
Gretel.aiGretel.ai
1
AppenAppen
1
DatabricksDatabricks
2
SnowflakeSnowflake
2
ConfluentConfluent
1
FivetranFivetran
1
dbt Labsdbt Labs
1
CollibraCollibra
1
Monte CarloMonte Carlo
1
BigIDBigID
1
PineconePinecone
1
QdrantQdrant
1
Zilliz / MilvusZilliz / Milvus
1
WeaviateWeaviate
1
ClickHouseClickHouse
1
ElasticsearchElasticsearch
1
OpenSearchOpenSearch
1
VespaVespa
1
Hugging FaceHugging Face
1
🛠️

Tech Stack Categories & Atomic Nodes

3 Tracks
3.1

3.1 Data Acquisition, Processing & Governance

Batch/stream ingestion, Lakehouse architecture (Iceberg/Delta), dbt modeling, metadata lineage, data contracts, and PII privacy governance.
Sub-Page
Lakehouse Batch/Stream Ingestion Pipelines
💰 $175K - $360K / year (Data Engineering) | ¥450K - ¥1.1M / year
High-throughput streaming/batch processing via Spark/Flink, Delta Lake/Iceberg/Hudi lakehouse formats, dbt SQL transformations, schema evolution, and multimodal OCR/PDF ingestion.
Benchmark:DatabricksDatabricksSnowflakeSnowflakeConfluentConfluentFivetranFivetrandbt Labsdbt Labs
Apache SparkApache FlinkApache IcebergDelta Lakedbt LabsSchema EvolutionMultimodal ParsingLakehouse
Data Governance, Lineage Tracking & Privacy Compliance
💰 $165K - $340K / year (Governance & Privacy) | ¥400K - ¥950K / year
Enterprise data catalogs, column-level lineage tracking, Data Contracts, automated PII redaction/masking, consent compliance, and dataset version snapshotting.
Benchmark:DatabricksDatabricksCollibraCollibraMonte CarloMonte CarloBigIDBigIDSnowflakeSnowflake
Data CatalogLineageData ContractsPII RedactionData PrivacyGreat ExpectationsConsent Governance
3.2

3.2 Training Data, Annotation & Synthetic Data

TB/PB-scale pretraining curation, MinHash LSH deduplication, benchmark decontamination, human-in-the-loop annotation, and verifiable synthetic data engines.
Sub-Page
PB-Scale Pre-training Curation & Decontamination
💰 $180K - $370K / year (ML Data Curation) | ¥480K - ¥1.15M / year
MinHash/LSH fuzzy deduplication, FastText quality/language scoring, HTML boilerplate removal, benchmark decontamination (preventing test set leakage), and multimodal pair filtering.
Benchmark:Scale AIScale AILabelboxLabelboxAppenAppenSurge AISurge AIHugging FaceHugging Face
MinHash LSHDeduplicationDecontaminationFastTextPII RemovalData Quality ScoringFineWeb
Synthetic Data Engines, Sandbox Verification & Provenance
💰 $195K - $410K / year (Synthetic Data R&D) | ¥500K - ¥1.3M / year
Teacher-model Evol-Instruct evolutionary scaling, code execution sandbox validation, rejection sampling, human-in-the-loop QA gates, and data provenance/licensing tracking.
Benchmark:Scale AIScale AIGretel.aiGretel.aiSnorkel AISnorkel AISurge AISurge AI
Synthetic DataEvol-InstructRejection SamplingExecution FeedbackData ProvenanceHuman-in-the-loopSelf-Instruct
3.3

3.3 Retrieval, Search & Vector Index Infrastructure

Inverted indexing, BM25 lexical search, HNSW/IVF-PQ vector engines, hybrid search (dense+sparse), index freshness, and ACL-aware retrieval infrastructure.
Sub-Page
Inverted Index, BM25 & Enterprise Search Engines
💰 $185K - $380K / year (Search & Relevance) | ¥450K - ¥1.1M / year
Distributed inverted indices in Elasticsearch/OpenSearch, BM25 scoring, tokenization, sharding, query understanding, ACL security filtering, and offline relevance evaluation (nDCG/MRR).
Benchmark:ElasticsearchElasticsearchOpenSearchOpenSearchVespaVespaClickHouseClickHouse
Inverted IndexBM25ElasticsearchOpenSearchVespaShardingACL RetrievalnDCG
Vector DB & High-Dimensional HNSW / IVF Indexing
💰 $190K - $395K / year (Vector Database Systems) | ¥480K - ¥1.2M / year
HNSW graph indexing & heuristic pruning, IVF-PQ product quantization, SIMD/GPU distance acceleration, hybrid scalar-vector filtering, and millisecond index freshness.
Benchmark:PineconePineconeZilliz / MilvusZilliz / MilvusQdrantQdrantWeaviateWeaviate
Vector DatabaseHNSW GraphIVF-PQEmbeddingPineconeMilvus / ZillizQdrantHybrid Retrieval
💼

Career Track Roles

13
💼Data Engineer
3
💼Analytics Engineer
1
💼Data Platform Engineer
2
💼Data Governance Engineer
1
💼ML Data Engineer
3
💼Data Operations Engineer
1
💼Annotation Operations Lead
1
💼Synthetic Data Engineer
2
💼Data Quality Engineer
2
💼Search Engineer
2
💼Relevance Engineer
1
💼Vector Database Engineer
1
💼Retrieval Engineer
2
🔄Sub-Domain Sequential Path (3 stages):
3.1

3.1 Data Acquisition, Processing & Governance

Open Sub-Page

Batch/stream ingestion, Lakehouse architecture (Iceberg/Delta), dbt modeling, metadata lineage, data contracts, and PII privacy governance.

Lakehouse Batch/Stream Ingestion Pipelines

High-throughput streaming/batch processing via Spark/Flink, Delta Lake/Iceberg/Hudi lakehouse formats, dbt SQL transformations, schema evolution, and multimodal OCR/PDF ingestion.

🏢 Companies
DatabricksDatabricksSnowflakeSnowflakeConfluentConfluentFivetranFivetrandbt Labsdbt Labs
🛠️ Tech Stack
Apache SparkApache FlinkApache IcebergDelta Lakedbt LabsSchema EvolutionMultimodal ParsingLakehouse
💼 Roles & Salary
Data Engineer、Analytics Engineer、Data Platform Engineer
💰 $175K - $360K / year (Data Engineering) | ¥450K - ¥1.1M / year

Data Governance, Lineage Tracking & Privacy Compliance

Enterprise data catalogs, column-level lineage tracking, Data Contracts, automated PII redaction/masking, consent compliance, and dataset version snapshotting.

🏢 Companies
DatabricksDatabricksCollibraCollibraMonte CarloMonte CarloBigIDBigIDSnowflakeSnowflake
🛠️ Tech Stack
Data CatalogLineageData ContractsPII RedactionData PrivacyGreat ExpectationsConsent Governance
💼 Roles & Salary
Data Governance Engineer、Data Platform Engineer、Data Quality Engineer
💰 $165K - $340K / year (Governance & Privacy) | ¥400K - ¥950K / year
3.2

3.2 Training Data, Annotation & Synthetic Data

Open Sub-Page

TB/PB-scale pretraining curation, MinHash LSH deduplication, benchmark decontamination, human-in-the-loop annotation, and verifiable synthetic data engines.

PB-Scale Pre-training Curation & Decontamination

MinHash/LSH fuzzy deduplication, FastText quality/language scoring, HTML boilerplate removal, benchmark decontamination (preventing test set leakage), and multimodal pair filtering.

🏢 Companies
Scale AIScale AILabelboxLabelboxAppenAppenSurge AISurge AIHugging FaceHugging Face
🛠️ Tech Stack
MinHash LSHDeduplicationDecontaminationFastTextPII RemovalData Quality ScoringFineWeb
💼 Roles & Salary
ML Data Engineer、Data Operations Engineer、Annotation Operations Lead
💰 $180K - $370K / year (ML Data Curation) | ¥480K - ¥1.15M / year

Synthetic Data Engines, Sandbox Verification & Provenance

Teacher-model Evol-Instruct evolutionary scaling, code execution sandbox validation, rejection sampling, human-in-the-loop QA gates, and data provenance/licensing tracking.

🏢 Companies
Scale AIScale AIGretel.aiGretel.aiSnorkel AISnorkel AISurge AISurge AI
🛠️ Tech Stack
Synthetic DataEvol-InstructRejection SamplingExecution FeedbackData ProvenanceHuman-in-the-loopSelf-Instruct
💼 Roles & Salary
Synthetic Data Engineer、ML Data Engineer、Data Quality Engineer
💰 $195K - $410K / year (Synthetic Data R&D) | ¥500K - ¥1.3M / year
3.3

3.3 Retrieval, Search & Vector Index Infrastructure

Open Sub-Page

Inverted indexing, BM25 lexical search, HNSW/IVF-PQ vector engines, hybrid search (dense+sparse), index freshness, and ACL-aware retrieval infrastructure.

Inverted Index, BM25 & Enterprise Search Engines

Distributed inverted indices in Elasticsearch/OpenSearch, BM25 scoring, tokenization, sharding, query understanding, ACL security filtering, and offline relevance evaluation (nDCG/MRR).

🏢 Companies
ElasticsearchElasticsearchOpenSearchOpenSearchVespaVespaClickHouseClickHouse
🛠️ Tech Stack
Inverted IndexBM25ElasticsearchOpenSearchVespaShardingACL RetrievalnDCG
💼 Roles & Salary
Search Engineer、Relevance Engineer、Retrieval Engineer
💰 $185K - $380K / year (Search & Relevance) | ¥450K - ¥1.1M / year

Vector DB & High-Dimensional HNSW / IVF Indexing

HNSW graph indexing & heuristic pruning, IVF-PQ product quantization, SIMD/GPU distance acceleration, hybrid scalar-vector filtering, and millisecond index freshness.

🏢 Companies
PineconePineconeZilliz / MilvusZilliz / MilvusQdrantQdrantWeaviateWeaviate
🛠️ Tech Stack
Vector DatabaseHNSW GraphIVF-PQEmbeddingPineconeMilvus / ZillizQdrantHybrid Retrieval
💼 Roles & Salary
Vector Database Engineer、Retrieval Engineer、Search Engineer
💰 $190K - $395K / year (Vector Database Systems) | ¥480K - ¥1.2M / year