- Implement the medallion architecture: design patterns
from the K&D Architect become working pipelines, tables, and
materialization schedules. Delta Lake tables designed with
partitioning, optimization, and evolution in mind.
- Own the workflow-orchestration framework: Databricks
Workflows, Delta Live Tables, retry policies, alerting routes, run
history, cost tags.
- Implement governance patterns: Unity Catalog structure,
sensitivity classification, access control, and audit for AI-facing
data assets. Design comes from the K&D Architect; day-to-day
implementation lives here.
- Build and operate retrieval-supporting infrastructure:
embedding pipelines, vector store maintenance, reindexing when
models upgrade, retrieval evaluation frameworks.
- Partner with Operations Capabilities on source-system
contracts: define what the AI Fabric consumes at the landing
zone — schema, cadence, SLA, quality thresholds. Own the
platform-side of that contract.
- Design and operate data quality: data-quality checks,
freshness monitoring, drift detection, and the alerting that
surfaces issues before they hit AI users.
- Support AI Developer velocity: when AI Developers deploy
into product teams, this role is the data engineer they turn to
when a use case needs specific data.
- Contribute to and consume the reference-pattern library:
reusable pipeline patterns, code templates, and standards live in
the shared platform layer. This role uses them, contributes new
ones, and evolves them as we learn.
What We Are Looking For
- Hands-on production data engineering: at least four
years designing, building, and operating production data pipelines.
Not a Databricks-course track record — real production experience
where you owned the pipeline through breakage, iteration, and
recovery.
- Databricks and Spark depth: Delta Lake, medallion
architecture, Delta Live Tables, Workflows, Databricks SQL, Unity
Catalog. Comfortable at the layer where code meets platform.
- Python and SQL fluency: PySpark, ETL patterns, and SQL
that runs at scale. Version control (Git), CI/CD for data
pipelines, and infrastructure-as-code (Terraform) as working
tools.
- AI-adjacent data engineering: real experience building
the data foundation for AI use cases — embedding pipelines, vector
stores, chunking strategies, retrieval evaluation. Not required to
be an ML researcher; required to have built the plumbing.
- Workflow orchestration: Databricks Workflows, or Airflow
in production. Retry semantics, dependency management, failure
handling — as working discipline, not concepts.
- Data quality and observability: Great Expectations,
Databricks data-quality monitors, or equivalent. Treats data
quality as a first-class engineering concern.
- Governance discipline: works with Unity Catalog
structures, understands sensitivity classification, and designs
pipelines with access control and audit in mind from the first
commit.
- AWS foundations: IAM, S3, KMS at the level needed to
work in a Databricks-on-AWS environment. Not required to be a cloud
architect; required to be productive.
- Communication: works productively with the K&D
Architect on design, AI Developer on integration, and Operations
Capabilities on contracts. Explains data-engineering trade-offs to
non-engineers.
- Education and experience: bachelor’s degree or
equivalent, plus at least four years of hands-on data-engineering
experience with meaningful exposure to AI or knowledge-management
use cases.
Nice to Have
- Prior experience with knowledge graphs (Neo4j or comparable),
entity resolution, or semantic data models.
- Experience with the modern data stack alongside
Databricks-native tooling.
- Familiarity with LLM-based extraction, chunking, and evaluation
frameworks.
- Background in research, academic, or mission-driven
institutional environments.
- Experience with cross-platform data engineering (Snowflake,
BigQuery) — cross-platform judgment is useful even when Databricks
is the primary tool.
What This Role Is Not
- Not a Data Platform Architect: design authority for the
Databricks-on-AWS platform sits with Platform Architects. This role
builds AI-facing data pipelines on that platform.
- Not the Knowledge & Data Architect: the K&D
Architect owns the design of the knowledge and retrieval layer;
this role implements against that design. Required to be a strong
hands-on builder, not an architect.
- Not an AI/ML engineer: the AI Developer owns model
access, orchestration, and application delivery. This role builds
the data layer AI Developers consume.
- Not a warehouse or BI data engineer: traditional
analytics warehouses and BI-facing pipelines belong to Operations
Capabilities. This role is AI-facing.
Practical Details
This role is hybrid, with three days per week in-person at HHMI’s
offices in Chevy Chase, Maryland. It reports to the Director of AI
Enablement. HHMI is not able to sponsor a visa for this position at
this time.
Physical Requirements
Remaining in a normal seated or standing position for extended
periods of time; reaching and grasping by extending hand(s) or
arm(s); dexterity to manipulate objects with fingers, for example
using a keyboard; communication skills using the spoken word;
ability to see and hear within normal parameters; ability to move
about workspace. The position requires mobility, including the
ability to move materials weighing up to several pounds (such as a
laptop computer or tablet).
Persons with disabilities may be able to perform the essential
duties of this position with reasonable accommodation. Requests for
reasonable accommodation will be evaluated on an individual
basis.
Please Note:
This job description sets forth the job’s principal duties,
responsibilities, and requirements; it should not be construed as
an exhaustive statement, however. Unless they begin with the word
“may,” the Essential Duties and Responsibilities described above
are “essential functions” of the job, as defined by the Americans
with Disabilities Act.
Compensation and Benefits
Our employees are compensated from a total rewards perspective in
many ways for their contributions to our mission, including
competitive pay, exceptional health benefits, retirement plans,
time off, and a range of recognition and wellness programs. Visit
our
Benefits at HHMI site to learn more.
Hiring Pay Range
$128,816.80 - $161,021.00
Pay Type:
Annual
The posted range reflects HHMI’s good faith estimate of the
anticipated hiring salary range for this role at the time of
posting. Actual hiring compensation is determined by a candidate’s
qualifications, experience, and internal equity.
HHMI is an Equal Opportunity
Employer
We use
E-Verify to confirm the identity and employment
eligibility of all new hires.