Senior Ontologist - Knowledge Graph & Identity
See all open roles at samba tv →
Some ghost-posting signals
- open for 167 days (90+ without a fill is a strong ghost signal)
- 67 open roles at this company in 30 days (mass-hiring blitz)
See your fit for this role and apply with a truthfully tailored résumé.
About the role
What You'll Do:
Ontology Design & Governance
Own the end-to-end design, development, and versioning of Samba TV's core ontologies in RDF/RDFS/OWL - defining entity classes, properties, hierarchies, and constraints that accurately model Samba's data domain at scale
Author and maintain SHACL shapes for post-load graph validation, consistency checking, and data quality enforcement
Define and document derived-attribute schemas - genre affinity, brand affinity, topic affinity, lifecycle signals, and viewing summaries - and own the logical definitions that govern how raw events become durable graph attributes
Establish ontology design standards, change management processes, and versioning practices; evaluate alignment with W3C standards and relevant industry schemas (Schema.org, EIDR, DDEX, W3C PROV)
Lead ontology design reviews with product, data engineering, and data science stakeholders - articulating trade-offs between expressivity, scalability, and query performance clearly
Event-to-Ontology Derivation
Define the aggregation and scoring logic that transforms raw TV viewership and web activity events into the durable affinities, summaries, and inferred signals that live in the graph
Co-own derivation pipeline design with data engineering - specifying transformation logic, intermediate schemas, and validation checkpoints for Databricks/Spark pipelines that feed the materialized graph substrate
Reason carefully about what belongs in the graph vs. what should remain virtualized in the data lake - balancing query performance against storage and refresh cost
Knowledge Graph Development & AI Integration
Build and maintain production-quality knowledge graph pipelines in Python and SPARQL - well-tested, documented, and scalable to Samba's data volumes
Design and implement entity resolution and record linkage pipelines that map real-world entities (content titles, devices, audiences, advertisers) to canonical knowledge graph nodes
Develop enrichment workflows that integrate third-party data sources (metadata providers, identity vendors, web sources) into Samba's knowledge graph in a consistent, governed way
Apply embedding-based and LLM-augmented approaches to ontology mapping, entity disambiguation, and semantic similarity problems
Support content and semantic embedding pipelines that feed into the vector store and underpin GraphRAG-based AI solutions
Cross-functional Collaboration & Mentorship
Partner with data engineering and platform teams to ensure the knowledge graph is integrated, queryable, and production-ready at scale
Collaborate with product to translate business requirements into ontological and graph data model decisions
Formally mentor Ontology Engineers and junior data scientists on semantic modeling, SHACL design patterns, and graph best practices
Lead internal technical talks and workshops on ontology, knowledge graph, and semantic web topics
Who You Are:
Must-Haves
5–8 years of hands-on experience in ontology engineering, semantic data modeling, or knowledge graph development - with a demonstrable track record of production ontologies at scale
Deep expertise in W3C semantic web standards: RDF, RDFS, OWL, SPARQL 1.1, and SHACL - with hands-on experience building and validating graph schemas in a production triplestore (Amazon Neptune, Stardog, GraphDB, Jena, or equivalent)
Strong Python - production-quality, well-tested code; comfortable building data pipelines and graph processing workflows
First-principles understanding of description logics, ontology design patterns, and the practical trade-offs between OWL expressivity and triplestore scalability
Hands-on experience with entity resolution, record linkage, or deduplication at scale - mapping messy, multi-source real-world data to clean ontological representations
Bachelor's degree required in Computer Science, Information Science, Computational Linguistics, Mathematics, or a related field; Master's or PhD strongly preferred
Strong communicator - able to defend ontological modeling decisions in design reviews and explain trade-offs to non-specialist stakeholders
Strongly Preferred
Hands-on experience with Amazon Neptune or Stardog - including data virtualization (Neptune Orion or Stardog Virtual Graphs) over data lake sources
Experience designing aggregation and derivation logic that converts raw behavioral event data into durable, graph-resident derived attributes
Domain knowledge in media, entertainment, or ad tech - TV viewership (ACR/STB), digital audience modeling (device graphs, identity resolution), or ad exposure data
Familiarity with industry content and identity schemas: EIDR, Schema.org VideoObject, DDEX, or equivalent
Experience with embedding models, vector databases (Milvus, Pinecone, Weaviate), and GraphRAG architectures (LangChain/LlamaIndex)
Familiarity with GNN-based approaches to knowledge graph reasoning or entity resolution a plus
Working knowledge of PySpark and Databricks for large-scale transformation pipelines
Stop applying to ghosts.
OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.
Do more with OyaPilot