Build scalable ETL/ELT pipelines in Databricks using PySpark, Spark SQL, Delta Lake, Delta Live Tables, Workflows, and Unity Catalog. Ingest batch and streaming data across medallion architecture layers, develop transformations and data quality rules, and implement orchestration, monitoring, alerting, and automation. The role requires cloud data engineering, data modeling, performance tuning, CI/CD, and strong knowledge of lakehouse architecture, governance, and distributed systems.
DATAECONOMY is one of the fastest-growing Data & Analytics company with global presence. We are well-differentiated and are known for our Thought leadership, out-of-the-box products, cutting-edge solutions, accelerators, innovative use cases, and cost-effective service offerings.
We offer products and solutions in Cloud, Data Engineering, Data Governance, AI/ML, DevOps and Blockchain to large corporates across the globe. Strategic Partners with AWS, Collibra, cloudera, neo4j, DataRobot, Global IDs, tableau, MuleSoft and Talend.
Senior/Lead Data Engineer — Databricks
Rutherford, NJ/ Jersey City, NJ
Full-time
- Build scalable, production-grade ETL/ELT pipelines using Databricks (PySpark, Spark SQL, Delta Live Tables, Workflows).
- Ingest structured, semi-structured, and streaming data into Bronze, Silver, and Gold layers.
- Develop optimized transformations, data quality rules, and reusable framework components.
- Implement best practices for job orchestration, monitoring, alerting, and automation.
- Hands-on experience: Spark, Delta Lake, Workflows, Unity Catalog.
- Strong SQL programming and performance tuning skills.
- Experience with cloud environments (AWS/Azure/GCP).
- Experience with modern data lakehouse concepts and distributed systems.
- Strong understanding of Lakeflow Connect, LSDP/Lakehouse, Medallion Architecture, Data Validations, Genie, and Agent Bricks/RAG use cases.
- Should be able to explain these concepts using real project examples and architecture decisions.
- Knowledge of medallion architecture, DLT and unity catalog within Databricks.
Requirements
- Strong Python (PySpark) and SQL programming
- Databricks — Spark, Delta Lake, Workflows, Unity Catalog
- ETL/ELT pipeline development — Medallion Architecture (Bronze/Silver/Gold)
- Delta Live Tables, Auto-Loader, Structured Streaming
- Data modeling — dimensional (star/snowflake), normalization/denormalization
- CI/CD, Git, job orchestration
- Cloud experience — AWS, Azure, or GCP
- 7–10+ years in data engineering
- Knowledge of medallion architecture, DLT and unity catalog within Databricks.
Nice-to-Have Skills
- Lakeflow Connect, LSDP/Lakehouse, Genie, Agent Bricks/RAG use cases
- Data governance, metadata management, Unity Catalog advanced features
- Airflow, dbt, or similar orchestration tools
- Data security, compliance, and access models
- Cost optimization and performance tuning in cloud environments
- Corporate/enterprise data warehousing background
Similar Jobs
Aerospace • Information Technology • Security • Cybersecurity • Defense
Lead database engineering for a federal species-data integration program in a hybrid AWS cloud environment. Responsibilities include harmonizing schemas, documenting source-to-target mappings and lineage, enforcing metadata and data-quality standards, implementing encryption and NIST-aligned security controls, building authoritative source registries, optimizing performance, and leading production migration, validation, monitoring, and operational documentation. The role requires Databricks or equivalent lakehouse expertise, ETL programming, federal governance knowledge, and database migration leadership.
Top Skills:
SparkAws GovcloudDatabricksDelta LakePythonScala
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead operational reliability and platform enablement for Databricks: build monitoring, CI/CD, deployment standards, compute and job policies, observability, runbooks, and governance to support secure, cost-aware, production data workloads across regulated environments. Mentor engineers and align platform with cloud/infrastructure and compliance requirements.
Top Skills:
Ci/CdDatabricksDatabricks Asset BundlesDatabricks WorkflowsDelta LakeInfrastructure-As-CodeService PrincipalsUnity CatalogVersion Control (Git)
Software
Design and build modern GCP lakehouse architectures, including Python and Spark pipelines, BigQuery ingestion, Kafka CDC, Delta Lake and Iceberg UniForm integrations, Delta Sharing endpoints, Snowflake and Databricks connectivity, data lineage, governed access controls, and semantic-layer alignment. Collaborate cross-functionally, troubleshoot full-stack issues, and deliver reusable data-sharing adapters and end-to-end data features.
Top Skills:
Apache IcebergSparkBigQueryChange Data Capture (Cdc)DatabricksDatabricks Unity CatalogDelta LakeDelta SharingGCPGoogle Cloud StorageIceberg UniformKafkaLookmlPythonSnowflakeSnowflake Horizon Catalog
What you need to know about the Colorado Tech Scene
With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.
Key Facts About Colorado Tech
- Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
- Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
- Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
- Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute



