Machine Learning / Data Engineer

Turing · Brazil; Colombia, Huila, Colombia; São Paulo, Brazil · Engineering

Posted 2026-10-07

Apply for this role →

Senior ML & Data Engineer — Data Quality & Sensitive Data Compliance

This is a full-time remote role based in Brazil or Colombia.

About the role

Enterprise data flows through our connectors, gets processed, and passes through a sanitization layer before anything downstream touches it. Two things have to be true at every step: the data is what we think it is, and no sensitive information — PII, PHI, company identifiable information (CII), or financial data — gets through. You'll own both. You'll do this primarily by building the machine learning that detects sensitive entities in text and image data and replaces them consistently at scale.

This is a hands-on IC engineering role with a QA mindset. You'll build the detection models, validation infrastructure, adversarial test sets, and audit processes that let us make strong claims about data quality and de-identification performance — and back them up with evidence. You'll work closely with a senior ML lead, with no client-facing responsibilities.

What you'll do

Data Quality

Run deep dives into enterprise data to assess quality: topic coherence across connectors, domain depth within connectors, completeness, and consistency

Design and automate validation suites for data pipelines — schema checks, completeness, drift detection, and reconciliation across raw → processed → sanitized stages

Surface and characterize quality issues in ways that engineering and product can act on

Sensitive data compliance (PII / PHI / CII / financial)

Design, train, and evaluate ML models (NER and other approaches) that detect sensitive entities across text and image-based documents such as scans, invoices, and presentations

Build replacement pipelines that substitute detected entities with coherent alternatives, so the same entity always maps to the same replacement across every file in a corpus and the data stays useful

Run these algorithms over large volumes of data to prepare it for downstream agentic task building

Build adversarial test sets for de-identification across all sensitive data classes: edge cases, obfuscated identifiers, multilingual entities, OCR noise, and formats designed to slip past detectors

Cover company identifiable information specifically — organization names and aliases, domains and email patterns, internal project and system names, org charts, vendor and partner relationships, contract terms, and any combination of details that could re-identify the source enterprise

Cover financial data — account and routing numbers, card numbers, revenue and pricing figures, transaction records, tax IDs, and financial statements

Measure and report de-identification performance by data class — entity-level precision and recall, leak rates, false-negative audits, and replacement consistency

Implement regression gates in CI/CD so no pipeline change ships without passing data quality and sensitive-data checks

Run sampling-based human-in-the-loop audits and maintain the audit trail as compliance evidence

Partner with engineering on root-cause analysis when inconsistencies or leaks are found, and drive fixes to closure

What we're looking for

About 4 to 5 years of hands-on machine learning experience, with ML as your primary background

Strong Python for ML development and data validation (pytest, Great Expectations, Pandera, or similar)

Solid SQL and experience validating data across pipeline stages

Familiarity with sensitive data categories and the relevant standards — HIPAA Safe Harbor for PHI, GDPR/LGPD for PII, PCI DSS for cardholder data, and confidentiality/NDA obligations for company information

Experience building and testing NER or other ML-based detection systems: building labeled eval sets, computing precision/recall, handling non-determinism

Understanding of re-identification risk — how seemingly innocuous details combine to reveal an organization or individual

Comfort with ambiguity and a fast-moving environment

A skeptical, detail-oriented approach — you assume things are broken until you've proven otherwise

Nice to have

Computer vision and OCR experience, especially building, scaling, and evaluating document pipelines for contracts, statements, invoices, presentations, and internal documents

Hands-on experience with financial or healthcare data, including the privacy requirements specific to those industries

Startup experience

Auditing LLM or VLM outputs

Synthetic sensitive-data generation (PII, PHI, company and financial records)

Familiarity with the GCP data stack (BigQuery, GCS, Cloud Run jobs) and CI/CD integration

Experience handling multi-tenant enterprise data with strict customer confidentiality requirements

Compliance reporting or working with auditors

Why this role matters

Our enterprise customers trust us with their data on the condition that it can never be traced back to them. Every downstream model, dashboard, and customer commitment depends on the data being clean and the sanitization layer being airtight. When you find a leak, you've prevented an incident. When you prove there isn't one, you've earned the trust that lets the rest of the company move fast.

Values

Apply for this role →

← Back to all jobs