Big Data Development Engineer Intern
Key Responsibilities:
Research & Benchmarking — Survey industry and open-source approaches in AI + data governance (metadata & lineage platforms, semantic layers, data quality frameworks, NL2SQL / NL2Metric); produce a best-practice proposal adapted to our tech stack.
Data Model Governance — Build AI-assisted model review capabilities: naming convention and layering (ODS/DWD/DWS/ADS) validation, duplicate and redundant model detection, lineage-based identification of unused / low-value / high-cost assets, and automated refactoring recommendations.
Metric Semantic Automation — Automatically extract and standardize metric definitions from SQL, lineage, and documentation; detect definition conflicts and redundant builds of the same metric across different reports; maintain a machine-readable semantic layer to support natural language metric queries.
Data Quality Automation — Auto-generate quality rules based on data profiling and lineage (replacing hand-written rules); perform anomaly detection on data volume, distribution, timeliness, and schema drift; conduct root cause analysis along lineage; implement automated alert grading and remediation recommendations (or auto-remediation).
MVP Delivery — Run at least one end-to-end implementation across the three areas: problem definition → solution design → prototype → deployment on a real data domain → quantified results (coverage, precision/recall of issue detection, manual effort saved) → iteration.
Documentation & Communication — Produce solution designs, evaluation methodologies, and results; present findings to platform and data stakeholders; deliver a reusable framework rather than one-off scripts.
Requirements:
Undergraduate or graduate student in Computer Science, Data Science, Statistics, or a related field.
Solid proficiency in SQL and Python. Understanding of data warehouse fundamentals — dimensional modeling, layered architecture, metadata, and lineage. Experience with Spark / Flink / Hive / StarRocks is a plus.
Hands-on experience with LLM application development: prompt engineering, RAG, Agent / tool-calling frameworks (e.g., LangChain, LlamaIndex, MCP), with the ability to evaluate whether an LLM system is actually effective.
Structured thinking: able to distill a vague governance pain point into a well-defined problem with quantifiable success criteria, and honestly articulate what the MVP validated and what it did not.
Self-driven and comfortable with ambiguity — this is an exploratory project with no predetermined answers.
Bonus: experience with DataHub / OpenMetadata / Atlas, dbt, Great Expectations / Deequ, or any metrics / semantic layer tooling.
Able to read technical materials in English; clear written communication skills.
Minimum 3-month internship commitment, 5 days per week on-site.
Fluent in Mandarin is required; fluent English is a plus.