Granica Logo

Granica

Senior Software Engineer – Foundational Data Systems for AI

Reposted 4 Days Ago
In-Office
Mountain View, CA
190K-250K Annually
Senior level
In-Office
Mountain View, CA
190K-250K Annually
Senior level
Design and implement foundational data systems for AI, focusing on efficiency and performance at scale. Collaborate on systems that optimize data handling and contribute to research advancements.
The summary above was generated by AI

Senior Software Engineer – Foundational Data Systems for AI

Location: Mountain View, CA — On-site

About Granica

Granica builds AI infrastructure for enterprises operating massive data environments.

Our platform helps data and engineering teams reduce storage and compute costs, improve performance and reliability, and prepare large datasets for analytics and AI.

Granica’s products include:

  • Crunch — continuous optimization for enterprise lakehouse data

  • Myelin — stateful infrastructure for long-running AI agents

  • Large Tabular Models — foundation models designed for enterprise tables

Together, we are building the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.

Granica has demonstrated approximately $200K in annualized value per petabyte and verified customer value within weeks.

About the Role

Granica is hiring a Senior Software Engineer to build foundational data systems for AI.

You will work on the core infrastructure behind Crunch, Granica’s continuous optimization product for enterprise lakehouse data. This includes systems for metadata management, table maintenance, file layout optimization, distributed compute, and workload-aware data reorganization across petabyte- and exabyte-scale environments.

This is a hands-on engineering role for someone who has deep systems experience and wants to build at the intersection of data lakes, distributed systems, storage engines, query performance, and AI infrastructure.

You will work closely with Granica Research, led by Prof. Andrea Montanari at Stanford, to translate ideas from information theory, probabilistic modeling, compression, and learning efficiency into production systems.

What You’ll Do
  • Build metadata and transaction systems for large-scale tabular datasets

  • Design systems that support time travel, schema evolution, partition evolution, and atomic consistency

  • Develop table-maintenance infrastructure for lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi

  • Optimize file layout, clustering, compaction, data skipping, indexing, and metadata pruning

  • Improve performance and cost efficiency across Spark, Flink, Trino, Presto, Databricks, Snowflake-adjacent, and cloud object storage environments

  • Work with columnar formats such as Parquet and ORC, including encoding, compression, layout, and read-path optimization

  • Build adaptive engines that learn from access patterns and workloads to reorganize data automatically

  • Develop distributed compute pipelines that scale predictively and remain reliable under failure

  • Debug performance bottlenecks across storage, metadata, query execution, network, and compute layers

  • Implement research-driven algorithms in compression, representation, layout optimization, and data efficiency

  • Contribute to open-source or publish research when appropriate

What We’re Looking For
  • Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure

  • Production experience with modern data lake or lakehouse technologies such as Spark, Iceberg, Delta Lake, Hudi, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems

  • Hands-on experience with columnar formats such as Parquet or ORC

  • Understanding of metadata-driven architectures, table formats, query planning, and physical data layout

  • Strong programming skills in Rust, Go, C++, or similar systems-oriented languages

  • Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency

  • A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end

Bonus
  • Experience contributing to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems

  • Experience with compaction, clustering, manifests, snapshots, metadata catalogs, schema evolution, partition evolution, or table garbage collection

  • Background in storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization

  • Research or open-source contributions in distributed systems, databases, storage, compression, indexing, or data processing

  • Interest in how physical data representation affects model training, inference, retrieval, and reasoning efficiency

Why Join Granica
  • Build foundational infrastructure for enterprise data and AI

  • Work on deep systems problems across lakehouse data, metadata, storage layout, distributed compute, and AI efficiency

  • Partner directly with Research, Product, Engineering, and company leadership

  • Help shape Crunch, Granica’s production data optimization platform for enterprise-scale lakehouse environments

  • Work with a small, high-caliber team solving high-value infrastructure problems at massive scale

  • Have direct influence on architecture, product direction, customer outcomes, and company growth

Compensation & Benefits
  • Competitive salary, meaningful equity, and performance bonus for top performers

  • 401(k) with company match, comprehensive health coverage, and unlimited PTO

  • Daily catered meals in our Mountain View office

  • Support for research, publication, and conference participation

At Granica, you'll help build the next generation of enterprise AI—from exabyte-scale data infrastructure, Large Tabular Models (LTMs), and stateful AI agents. Together, we're creating the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.

 

Similar Jobs at Granica

3 Days Ago
Hybrid
160K-240K Annually
Senior level
160K-240K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Build and optimize distributed compute infrastructure for large-scale analytical and AI workloads. Responsibilities include improving query execution, scheduling, resource allocation, reliability, workload routing, and compute efficiency across Spark and related systems. The role involves debugging performance bottlenecks, optimizing joins, scans, shuffles, caching, partitioning, and memory usage, and working with lakehouse formats and cloud object storage. Candidates will implement workload optimization algorithms and may contribute to open source or research.
Top Skills: Adaptive Query ExecutionAmazon EmrAmazon S3Apache FlinkApache HiveApache HudiApache IcebergSparkAws GlueAzure Data Lake StorageC++CatalystDatabricksDatafusionDelta LakeDuckdbGoGoogle Cloud StorageJavaOrcParquetPrestoRustScalaSnowflakeSpark SqlTrinoVelox
3 Days Ago
In-Office
160K-240K Annually
Senior level
160K-240K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Build foundational lakehouse infrastructure for exabyte-scale AI data environments. Responsibilities include metadata and transaction systems, table maintenance, schema and partition evolution, snapshot isolation, compaction, clustering, file-layout optimization, object-store performance, columnar-format optimization, and query performance across major lakehouse engines. The role also involves debugging distributed systems, implementing compression and data-efficiency algorithms, and contributing to open-source or research efforts.
Top Skills: Amazon S3Apache FlinkApache HudiApache IcebergSparkAzure Data Lake StorageC++DatabricksDelta LakeGoGoogle Cloud StorageHive MetastoreJavaOrcParquetPrestoRustScalaSnowflakeTrinoUnity Catalog
5 Days Ago
In-Office or Remote
Senior level
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
The role involves designing cloud infrastructure, managing production Kubernetes clusters, optimizing CI/CD pipelines, enhancing developer experience, and ensuring reliable AI workloads. Candidates should have extensive experience in infrastructure and distributed systems engineering with strong coding skills and cloud expertise.
Top Skills: AWSAzureDatadogDockerElkGCPGoGrafanaJavaKubernetesPrometheusPythonTerraform

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account