Senior Data Scientist
Job Title : Senior Data Scientist
Summary:
The Senior Data Scientist – Knowledge Representation is a core technical role on the Healthcare Supply Chain (HCSC) Knowledge Representation (KR) team. This role brings machine learning, data science, and AI capabilities directly in service of knowledge representation problems: entity resolution, knowledge graphs, embedding-based semantic alignment, LLM-assisted knowledge extraction, and validation of ontological claims against data.
Our platform connects hospitals, distributors, GPOs, manufacturers, and regulators, enabling transactional execution, clinical data alignment, and analytics optimization across organizational boundaries. Each party encodes its own implicit knowledge about the world in its schemas, identifiers, and workflows. The KR team's job is to make that implicit knowledge explicit, alignable, and trustworthy.
Essential Duties:
Design and implement knowledge representation solutions that balance performance and costs: using a variety of data stores (graph, document, tabular, vector).
Design, build, and maintain entity resolution systems.
Develop and own embedding-based and hybrid semantic similarity solutions.
Build and operate LLM-assisted knowledge extraction pipelines that identify ontological candidates from unstructured and semi-structured HCSC data sources.
Design and implement uncertainty quantification frameworks for KR outputs: confidence scoring for entity resolution decisions, calibration of alignment scores, and propagation of data quality signals through KR pipelines so that downstream consumers understand the reliability of what they receive.
Partner with data quality engineers to establish bidirectional feedback channels between KR pipeline failures and systematic data quality issues.
Collaborate with internal and external stakeholders.
Stay current with developments in AI, ML and Data Science including knowledge graph embeddings, neuro-symbolic AI, and AI reasoning.
Competencies:
Strong ML engineering fundamentals: experience designing, training, evaluating, and operating models in production.
Genuine engagement with knowledge representation concepts: sufficient familiarity with OWL, description logics, and ontological commitments to reason about what an ML model's outputs actually assert.
Understanding and expertise in embedding methods for knowledge graphs and text with practical experience applying them to similarity, matching, and alignment tasks, and understanding of their failure modes in sparse or structurally heterogeneous domains.
Practical LLM engineering skills: prompt engineering for structured extraction, fine-tuning or adapter methods for domain adaptation, retrieval-augmented generation for KR tasks, and rigorous evaluation of LLM outputs.
Proficiency in knowledge graph technologies (RDF, SPARQL, property graphs) at a working level sufficient to query, evaluate, and annotate KR structures.
Excellent cross-disciplinary communication: able to explain statistical model behavior and uncertainty to stakeholders.
Comfort operating at the boundary of statistical and formal methods.
Requires minimal supervision on ML and data science work within the KR domain.
Required Qualifications and Skills:
Greater than 4 years of applied ML and data science experience, with at least some portion of that work directly involving entity resolution, semantic matching, knowledge graph construction, or related knowledge representation problems.
Hands-on experience with LLM-assisted knowledge extraction: prompt design, structured output parsing, domain fine-tuning or few-shot adaptation.
Working knowledge of RDF and SPARQL.
Demonstrated ability to design and implement model evaluation frameworks.
Strong programming skills for ML pipeline development, data analysis, and tooling; proficiency with the standard ML stack (e.g. PyTorch or equivalent, HuggingFace, scikit-learn) and with graph data tooling (RDFLib, NetworkX, or equivalent).
Experience working in multi-disciplinary settings where ML outputs feed into formal systems or decision-making processes.
Preferred Qualifications and Skills:
Bachelor's or advanced degree in Computer Science, Statistics, Computational Linguistics, Information Science, or a related discipline; graduate work in knowledge representation, NLP, or information extraction is a strong plus.
Experience with neurosymbolic AI approaches: methods that combine learned representations with formal reasoning, constraint satisfaction, or logic-based inference — particularly in the context of knowledge graph completion, ontology alignment, or structured prediction tasks.
Healthcare supply chain domain knowledge.
Familiarity with OWL 2 and description logics at a reading level: not required to author axioms independently, but able to read OWL and understand what a reasoner computes.
Experience with graph database platforms at production scale.