Senior Software Engineer, Data Engineering
See all open roles at valgenesis, inc →
Some ghost-posting signals
- open for 101 days (90+ without a fill is a strong ghost signal)
- 46 open roles at this company in 30 days (mass-hiring blitz)
- no salary disclosed (correlates with ghost postings)
See your fit for this role and apply with a truthfully tailored résumé.
About the role
Responsibilities:
Design, develop, and maintain data ingestion, transformation, and orchestration pipelines (batch and real-time).
Build and optimize data Lakehouse architectures using Azure Synapse, Delta Lake, or similar frameworks.
Integrate and manage structured and unstructured data sources (SQL/NoSQL, files, documents, IoT streams).
Develop and operationalize ETL/ELT pipelines using Azure Data Factory, Databricks, or Apache Spark.
Collaborate with Data Scientists to prepare and serve ML-ready datasets for model training and inference.
Implement data quality, lineage, and governance frameworks across pipelines and storage layers.
Work with BI tools (Power BI, Superset, Tableau) to enable self-service analytics for business teams.
Deploy and maintain data APIs and ML models in production using Azure ML, Kubernetes, and CI/CD pipelines.
Ensure scalability, performance, and observability of data workflows through effective monitoring and automation.
Collaborate cross-functionally with engineering, product, and business teams to translate insights into action.
Requirements:
Experience: 4 to 8 years in Data Engineering or equivalent roles.
Programming: Strong in Python, SQL, and at least one compiled language (C#, Java, or Scala).
Databases: Experience with relational (SQL Server, PostgreSQL, MySQL) and NoSQL (MongoDB, Cosmos DB) systems.
Data Platforms: Hands-on experience with Azure Data Lake, Databricks.
ETL/ELT Tools: Azure Data Factory, Apache Airflow, or dbt .
Messaging & Streaming: Kafka, Event Hubs, or Service Bus for real-time data processing.
AI/ML Exposure: Familiarity with ML frameworks (TensorFlow, PyTorch ) and MLOps concepts.
Visualization & Analytics: Power BI, Apache Superset, or Tableau.
Cloud & DevOps: Azure, Docker, Kubernetes, GitHub Actions/Azure DevOps.
Best Practices: Solid understanding of data modeling, version control, and CI/CD for data systems.
Preferred Skills (Nice to Have):
Experience in knowledge graph or semantic search solutions.
Understanding of LLM-based data retrieval (RAG) patterns.
Exposure to data mesh, data fabric, or domain-oriented data architecture.
Familiarity with MLflow , Delta Live Tables, or DataBricks Unity Catalog.
Soft Skills:
Strong analytical and problem-solving ability.
Excellent communication and collaboration skills.
Attention to detail and ability to work with large, complex datasets.
Creativity and ability to automate repetitive workflows.
Passion for continuous learning and innovation in data and AI technologies.
Stop applying to ghosts.
OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.
Do more with OyaPilot