Analytics Engineer, Life Sciences Delivery Operations
See all open roles at arcadia →
Likely real
- 17 open roles at this company in 30 days (mass-hiring blitz)
See your fit for this role and apply with a truthfully tailored résumé.
About the role
What You'll Be Doing
RWD DATA PIPELINE ENGINEERING
Author and maintain dbt models and PySpark transformation jobs, replacing ad-hoc Snowflake scripts with governed, version-controlled, tested code
Design and implement delivery endpoint configurations as code-customer, delivery target (Snowflake, S3), cadence, cohort filters, incremental and full historical refresh methods
Write production-grade Python and PySpark for data transformation, validation automation, and delivery pipeline components, including customer-specific data models and schema validation logic
Configure and maintain AWS S3 delivery paths, Apache Iceberg table structures, and file staging patterns for partner data delivery
Partner with platform engineering to build and extend Argo Workflows orchestration for automated delivery execution, eliminating the manual monthly Snowflake runbook
Implement and maintain HIPAA de-identification compliance rules in pipeline code in accordance with ED certificates; coordinate certification updates when new data elements or rule changes affect certified products
DELIVERY OPERATIONS & DATA QUALITY
Coordinate and execute monthly RWD deliveries across all active channel partners: delivery job execution, manifest generation and validation, tokenization workflows, and QC
Define and monitor delivery quality metrics: pipeline health, data completeness, referential integrity, refresh SLA tracking, and minimum volume thresholds; manage on-time delivery against a >=95% target
Validate dbt model outputs and pipeline changes against expected schema and counts; define acceptance criteria and execute UAT before changes reach production or channel partners
DATA INVESTIGATIONS & PARTNER SUPPORT
Own the channel partner data inquiry queue-triage, investigate, resolve, and communicate on data questions and discrepancies; you are the primary research contact for channel partners
Conduct root cause analysis on anomalies in RWD, claims, and clinical feeds, distinguishing source-level issues from transform-layer failures, and communicate findings clearly in writing
Build and maintain repeatable query libraries, data dictionaries, and end-to-end pipeline documentation in Confluence, reducing one-off analytical effort and improving institutional knowledge
Build out a knowledge repository of learnings from all data inquiries, and partner with VP to provide a data FAQ for LS customers
Contribute to data quality governance: build scorecards and trend analyses that equip leadership with evidence-based positions for engineering and product discussions
ENGINEERING PRACTICES & COLLABORATION
Follow SDLC best practices: author requirements, write test plans, manage releases, and maintain operating documentation in Confluence
Manage engineering work through Jira: clear ticket authoring, acceptance criteria, dependency tracking, and proactive status communication
Apply CI/CD practices via GitHub: branch management, PR-based review workflows, dbt model testing, and version control discipline-the version in production always matches what's in Git
Partner closely with Arcadia's platform engineering team on pipeline architecture, table design, and data contracts
Leverage AI tools (including Claude Code) to accelerate development, automate documentation, generate and verify code, and improve operational throughout
TECHNOLOGIES
Pipeline & transformation: dbt-Spark, PySpark, Python, SQL
Orchestration: Argo Workflows, Kubernetes
Storage: AWS S3, Apache Iceberg, AWS Glacier
Warehousing: Snowflake (delivery target, Iceberg, analytics)
Person De-identification and Tokenization Processes
Source control & CI/CD: Git / GitHub, PR-based review workflows
Workflow & observability: Jira, Confluence, etc.
AI tooling: Claude, ChatGPT
What You'll Bring
Education
Bachelor's or Master's degree in Computer Science, Data Science, Statistics, or a related field (or equivalent professional experience)
Experience
5+ years of hands-on data engineering experience (production pipelines, dbt, Spark/PySpark, cloud data infrastructure) AND 5+ years of direct experience with life sciences RWD data (claims, EHR, clinical); these disciplines can overlap-5 years total is sufficient if you bring meaningful depth in both
Production-grade SQL proficiency in Snowflake or a comparable columnar warehouse: complex joins, CTEs, window functions, incremental patterns – you write this fluently
Python and/or PySpark for data transformation: you have written and debugged production Spark jobs, not just automation scripts
dbt: hands-on experience authoring models, tests, macros, and yml documentation; familiarity with incremental strategies and model validation
AWS S3: practical experience with file staging, delivery paths, bucket structure, and lifecycle management in a data engineering context
HIPAA de-identification: working knowledge of Safe Harbor requirements and how they are applied in data pipelines before data leaves your custody
SDLC fundamentals: you write requirements, author test plans, manage releases, and document your work – this is not new to you
CI/CD and source control: Git/GitHub, PR-based review workflows, branching strategies
External customer experience: you have led (or actively presented within) technical data discussions with partner analytics or science teams and can communicate complex data concepts clearly in writing and verbally
Self-starter who operates independently in ambiguous, high-growth environments and a natural collaborator when the work calls for it
Skills
Strong analytical judgment – you can look at a distribution and know when something is wrong
Clear communicator – able to translate technical pipeline findings for partners and non-technical stakeholders
Builder's mindset – you don't just answer the question in front of you, you eliminate the conditions that caused it
Genuine curiosity – AI tooling and motivated to apply it to your own workflows, not just in principle
Multi-tasking ability – manage multiple parallel workstreams, triage competing priorities, and communicate proactively on risk
Would Love For You To Have
Experience with Argo Workflows or a comparable orchestration platform
Familiarity with patient tokenization or record linkage technologies and how they integrate into RWD delivery workflows
Experience with HIPAA de-identification certification process and implementation workflow
Healthcare data standards: ICD-10, CPT, NDC, LOINC, NPI
Experience working at a healthcare data vendor, RWD aggregator, or analytics platform company
Comfortable operating in an AI-first environment, using Claude or similar tools to build, verify, and accelerate day-to-day workflows
Demonstrated interest in people leadership-you'd be excited to mentor and eventually grow a small team as the LS organization scales
What You'll Get
Be at the center of a high-stakes, high-impact engineering RWD delivery pipeline you help create will define how Arcadia delivers RWD to life science partners at scale
Become the definitive internal expert on one of the most complex and valuable real-world healthcare datasets in the market, with the autonomy to shape how it is engineered, measured, and delivered
Be on the front lines of AI adoption-use cutting-edge tools to accelerate your work and shape how the team operates in an AI-first environment
Flexible, fully remote work environment, with resources and support to do your best work
Exposure to senior leaders across the entire life science and corporate engineering teams
A clear path to grow into a player/manager role as Arcadia's life sciences delivery team scales
Become a member of the talented, energized, diverse, and purpose-driven Arcadian community
Stop applying to ghosts.
OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.
Do more with OyaPilot