← All verified jobs

Member of Technical Staff, Distributed Systems

psi polska sp. z o.o. BostonRemoteFullTime

See all open roles at psi polska sp. z o.o.

Ghost-risk verdict

Strong ghost-posting signals

  • open for 95 days (90+ without a fill is a strong ghost signal)
  • reposted 1× (reposts correlate with ghost postings)
  • removed from the board and reposted at least once
  • 18 open roles at this company in 30 days (mass-hiring blitz)
  • no salary disclosed (correlates with ghost postings)

How we score ghost risk →

See your fit for this role and apply with a truthfully tailored résumé.

See my fit, free

About the role

Overview

Physical Superintelligence is a startup with roots at Google, NVIDIA, Harvard, Meta, MIT, Oxford, Johns Hopkins, Cambridge, and the Perimeter Institute building AI systems to discover new physics at scale. We are seeking engineers to build platform infrastructure at the intersection of computational science, AI systems, and software engineering.

Our mission is to discover and commercialize transformative physics breakthroughs at scale with artificial superintelligence, safely, verifiably, and for broad public benefit.

The last century's golden age of physics gave us transistors, lasers, and nuclear energy. We believe artificial superintelligence will unlock the next one. We're creating the infrastructure to industrialize scientific discovery and usher in this new era.

We have one product: new physics, at scale.

We are seeking a Member of Technical Staff, Distributed Systems to build the execution layer everything at PSI runs on: the systems that decide what runs where, keep it alive through failure, and make every run replayable and priced. A research campaign is a great many jobs that have to survive machines dying and still add up to a result somebody can reproduce a year from now. That is this layer's problem.

Role and Responsibilities

Design the runtime primitives researchers and engineers compose into agentic workflows: sequential pipelines, tree-search agents, and whatever pattern the science asks for next. Treat the platform as a library product. Clear layers, explicit API contracts, surfaces other engineers extend instead of fork.

Build the durable execution layer. A researcher's request becomes work that finds the right cluster, runs there, survives partial failure, and comes back replayable. Correctness under retries and replays. Idempotency as a design default. Admission and placement when the cluster is full, and gang scheduling for the jobs that need it.

Make every run priceable and replayable end to end. Every call traced, cost recorded at call time, one correlation ID from request to result.

Operate what you build. Set SLOs and meet them, build the instrumentation, plan capacity with the infrastructure team, and take your turn in incident response for the systems you own.

What We're Looking For

Five or more years building and operating distributed systems in production at companies known for engineering rigor, on major cloud platforms with Kubernetes, Slurm, or comparable orchestration. You have written code that paying customers or internal teams depend on every day.

You have implemented or substantially extended a durable scheduler, workflow engine, dataflow system, or agent runtime. You know the bugs that come out of retries, replays, and non-deterministic execution. You have the failure stories that prove you operated it, not only wrote it.

You have shipped a library or internal framework other engineers extend rather than only consume. API ergonomics, composability, and backward-compatible evolution are first-order concerns for you.

You have made a build-versus-adopt call on an orchestration engine and can argue both sides. You favor the boring, durable answer. You can explain a workflow-engine tradeoff in two minutes.

Nice to Have

Hands-on with a durable workflow system at scale: Temporal, Cadence, Step Functions, or Argo Workflows.

Data-infrastructure depth alongside the systems work: content addressing, catalogs, lineage. Our execution and data work sit next to each other, and the engineer strong in both is the one we most want to meet.

Systems where tracing and metering were product requirements, not afterthoughts; production observability on OpenTelemetry or comparable.

Background in scientific computing, HPC environments, or research infrastructure.

How We Work

We hold a high technical bar and give people full ownership of their work, from spec to ship to on-call. We write contracts before logic, test against real systems instead of mocks, and favor simple designs that ship over clever ones that do not. Our development process is AI-native: we work with agentic coding tools daily, write specs that are legible to humans and agents alike, and lead with leverage.

Location and Compensation

This role is based in Boston. We will consider remote candidates on a case-by-case basis. We offer competitive compensation including salary, benefits, and meaningful early-stage equity. We evaluate on technical breadth, systems thinking, scientific curiosity, and shipping velocity. We are an equal opportunity employer and value diverse perspectives in building platforms for AI-driven discovery.

Stop applying to ghosts.

OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.