Troveo AI Logo

Troveo AI

Data Engineer

Reposted 2 Hours Ago
Remote
Hiring Remotely in USA
100K-140K Annually
Senior level
Remote
Hiring Remotely in USA
100K-140K Annually
Senior level
Design, build, and maintain scalable ELT/ETL and streaming data pipelines into cloud warehouses. Implement transformations and dimensional models (dbt), ensure data quality and observability, write SQL for analysis, support BI/dashboarding, monitor SLAs, troubleshoot incidents, participate in on-call rotations, and collaborate cross-functionally to deliver analytics-ready datasets.
The summary above was generated by AI

About Troveo

Troveo builds the data platform that AI labs and model builders need to train the next generation of models. We have created the world's largest licensed platform of scarce, proprietary data for AI, spanning video, audio, text, and business workflows.

Troveo indexes, enriches, and packages this high-quality data into formats ready for training, fine-tuning, evaluation, and agentic use cases. Backed by top investors, we’re a small, high-impact team solving one of the biggest bottlenecks in AI development.

Role Overview

We are seeking a versatile, hands-on Data Engineer to build and maintain a scalable analytics data warehouse while contributing to data modeling, performing data analysis, and ensuring the reliable delivery of data to downstream teams and systems. This hybrid role combines core data engineering responsibilities with data modeling, analytics, and operational support. You will own the full analytics data lifecycle, from ingestion and transformation to modeling, quality assurance, and timely delivery, while partnering closely with software engineers and business stakeholders.

Key Responsibilities

Data Pipeline Engineering

  • Design, build, and maintain ELT/ETL data pipelines (batch and streaming), optimizing for performance, reliability, scalability, and cost.

  • Design and implement conceptual, logical, and physical data models (including dimensional modeling, star/snowflake schemas).

  • Build and maintain transformation layers using modern tools (e.g., dbt) to create clean, well-documented, analytics-ready datasets.

  • Apply data modeling best practices, versioning, testing, and documentation to ensure consistency and reusability.

Data Analysis & Reporting Support

  • Write optimal SQL queries for data exploration, ad-hoc analysis, and troubleshooting.

  • Support the creation of reports, dashboards, and self-service analytics assets in collaboration with data analysts and business teams.

  • Translate business questions into data requirements and deliver actionable insights or datasets.

Operational Support & Data Deliveries

  • Monitor data pipelines and data delivery processes to ensure SLAs for timeliness, freshness, and accuracy are consistently met.

  • Proactively identify, troubleshoot, and resolve data issues impacting downstream consumers or business operations.

  • Manage incidents related to data availability and quality; participate in on-call rotations as needed.

  • Implement data quality checks, observability, and alerting to maintain high reliability of data deliveries.

  • Automate operational tasks and continuously improve data delivery processes.

Collaboration & Best Practices

  • Work cross-functionally with analysts, data scientists, engineers, and business stakeholders to understand data needs and deliver solutions.

  • Document data pipelines, models, lineage, and processes.

  • Contribute to data governance, security, and best practices across the data platform.

Requirements

  • 7+ years of professional experience in data engineering or a closely related role (analytics engineering experience is highly relevant).

  • Strong proficiency in SQL and Python.

  • Hands-on experience building and maintaining data pipelines and working with cloud data platforms/warehouses. (Snowflake, BigQuery, Redshift, Databricks, etc.).

  • Experience with data orchestration tools (Apache Airflow, Dagster, Prefect, or similar).

  • Solid understanding of data modeling techniques and dimensional modeling.

  • Experience performing data analysis and working with BI/visualization tools (Looker, Tableau, Power BI, or similar).

  • Proven ability to troubleshoot data issues and support operational reliability/SLAs.

  • Strong communication skills and ability to collaborate with both technical and non-technical stakeholders.

Bonus Points

  • Experience with DBT for data transformation and modeling.

  • Knowledge of data observability/monitoring tools.

  • Experience with real-time/streaming data technologies (Kafka, Flink, etc.).

  • Familiarity with CI/CD practices for data pipelines.

  • Experience in data quality frameworks and governance.

  • Bachelor’s degree in Computer Science, Engineering, or a related quantitative field (or equivalent practical experience).

Compensation

  • Base Salary: $100,000 – $140,000 (depending on experience and location)

  • Equity: Competitive equity package in a well-funded AI startup with significant upside

Compensation is location-adjusted for cost of living. We are open to candidates in California, New York, and select other states.

What We Offer

  • Comprehensive Health Benefits: Medical, dental, and vision coverage (100% employer-paid for employees)

  • Flexible PTO & Paid Holidays: Unlimited PTO with encouragement to actually use it

  • Remote First Policy: Work from anywhere in the US (with occasional team offsites)

  • Learning & Growth: Annual learning stipend, access to top conferences, and direct mentorship from experienced founders

  • Equity Ownership: Competitive equity package with clear growth potential as we scale

  • Modern Tech Stack & Tools: Budget for the best equipment and software

  • Strong Culture: High-trust, low-ego environment focused on impact, transparency, and work-life balance. We believe great work happens when people are supported, challenged, and given ownership.

Equal Opportunity Employer

Troveo is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Similar Jobs

2 Days Ago
Remote or Hybrid
79K-135K Annually
Senior level
79K-135K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, build, and maintain large-scale ETL pipelines and data warehouses (Redshift) with strong data modeling, governance, and quality controls. Optimize systems for performance and reliability, integrate and deploy AI/ML models, troubleshoot and support Redshift, collaborate with data scientists and stakeholders, and produce technical documentation and knowledge-base articles.
Top Skills: Amazon RedshiftApache AirflowApache KafkaAWSAws GlueFivetranJavaNoSQLOraclePysparkPythonPyTorchScalaScikit-LearnSparkSQLSQL ServerTensorFlow
14 Days Ago
Remote or Hybrid
United States
70K-120K Annually
Mid level
70K-120K Annually
Mid level
Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Build, model, and maintain scalable BigQuery-based data solutions on GCP. Implement performant data models, storage partitioning/clustering, ETL improvements, data validation, and documentation. Collaborate with architects, data scientists, and engineers to deliver governed, high-quality data for reporting and AI initiatives.
Top Skills: Ansi SqlBigQueryBigquery SqlConfluenceGCPJIRAPythonSQL
7 Days Ago
Easy Apply
Remote
United States
Easy Apply
Senior level
Senior level
Enterprise Web • Mobile • Professional Services • Software
Design, build, and own scalable data pipelines and evaluation systems that power production AI features and internal analytics. Ensure data quality across ingestion, modeling, and reporting, collaborate with ML and analytics teams, deploy and monitor ML systems, and establish standards for data work and evaluation.
Top Skills: AirflowAWSDagsterGCPPostgresPythonSnowflake

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account