← All verified jobs

Senior AI Engineer

Ro New York, NY

See all open roles at Ro

Ghost-risk verdict

Some ghost-posting signals

  • open for 40 days (30+ days starts to look stale)
  • 56 open roles at this company in 30 days (mass-hiring blitz)
  • no salary disclosed (correlates with ghost postings)

How we score ghost risk →

See your fit for this role and apply with a truthfully tailored résumé.

See my fit, free

About the role

What You'll Do

Design, build, and ship production LLM powered features end-to-end.

Build the AI application layer, including prompt orchestration, tool calling, retrieval (RAG), embeddings, agent workflows, and structured outputs.

Develop evaluation infrastructure, including regression suites, LLM-as-a-judge evaluations, synthetic datasets, human review workflows, and quality metrics that catch regressions before users do.

Design observability systems that measure accuracy, latency, cost, hallucination rates, and model behavior in production.

Build safety and reliability guardrails that ensure AI systems meet quality, compliance, and operational requirements.

Evaluate and integrate third-party tools, including orchestration frameworks, evaluation platforms, vector databases, and model providers; and make pragmatic build-versus-buy decisions.

Raise engineering standards across the team through technical leadership, code reviews, mentorship, and thoughtful system design.

Sit with EHR operators, Patient Advocates, and clinical teams to understand the workflows you're automating.

Prototype quickly, ship early, measure outcomes, and iterate based on real-world usage.

What You’ll Bring to the Team

5+ years building production software, with at least 1–2 years building and deploying LLM-powered applications in production (or equivalent depth through substantial side projects or open-source work).

Strong Python engineer with experience building backend services and integrating modern AI APIs and frameworks.

Comfortable using AI-assisted development tools throughout the software development lifecycle.

You've shipped AI features that served real users, and you can clearly articulate how you measured quality, reliability, latency, cost, and business impact.

You have strong product instincts and thrive in ambiguous environments, translating operational problems into technical solutions without relying on detailed specifications.

Bonus: Experience with orchestration frameworks (LangChain, LangGraph, AgentCore), RAG architectures, vector databases, prompt/version management, and evaluation or observability platforms (e.g., LangSmith, Braintrust, Arize Phoenix).

Stop applying to ghosts.

OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.