Design and own the verification engine that automates product truth verdicts: build confidence-scoring pipelines, adversarially robust models, a human-graded golden exam, and arbitration rulebooks. Raise machine-settled verdicts from 40% to 80% while maintaining 95%+ accuracy, embed eval-driven development, and scale solutions to millions of claims in production.
You build the judge that decides what we publish as true - for millions of shopping claims, against merchants, bots, and models that all want to fool it.
Product.ai is the verified truth layer for shopping: the intelligence that tells you what is actually true about a product, including when not to buy. SimplyCodes is its first proof at scale: the code verification service - codes that actually work, proven by robots running real checkouts - at around $22M in revenue and roughly 60% margins. Founder-owned and profitable since 2009. No outside investors. No board. Fewer than twenty operators outbuilding companies 10x our size.
Strong people find us and keep finding us - they apply over months and years, because the field moves fast and the exact profile we need moves with it.
Why This Role Exists
Everything we sell rests on one question: is the claim true? Machine verdicts answer it - they decide what we publish about a product, a price, a code. Today machines settle a minority of those verdicts and humans settle the rest. The constraint on growth is not more claims; it is whether the machine verdicts stay right as the volume climbs. You own that: the eval science behind every automated verdict - the judges, the thresholds, and the exams that grade them. You work directly with the founder.
The target is the large majority of verdicts settled by machine, holding a high accuracy floor against a held-out, human-graded exam. You co-own the arbitration rulebooks - the rules that settle a contested verdict - with our truth scientist, so the truth an entire company ships through never rests on one person's judgment. It is also the most durable engineering seat in the building: every model generation absorbs another layer of surface skill, but grading a verdict no human can check faster than the machine is an open frontier that compounds for years.
The System You'll Need to Model
If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.
What You Will Own
You will use the craft you already own - judge design, golden sets, calibration, regression corpora, statistical rigor - and grow into the layer above it: eval-gated agent orchestration, adversarial verification at commerce scale, and adjudication design for judgments no single grader can settle.
Who You Are
You interrogate every green checkmark. A passing eval is a claim, and claims get challenged - you ask what the test could not have caught before you ask what it confirmed. You form working models of complex systems on your own, notice where your model is wrong, and update fast. You write clearly, because clear writing is evidence of clear thought.
You treat agents as leverage you verify, never as an oracle you trust. You can do this job by hand and prove it - hand-grade a judge, hand-build a corpus, hand-verify a claim set - and that mastery is exactly what lets you trust, or reject, the verdict an agent hands back. The expensive thing here is a redo cycle, never the compute.
You have probably built an eval harness another team came to depend on, a regression corpus that caught a real failure before users did, an LLM-judge pipeline where you measured the judge's bias instead of trusting it, or a golden set for a system whose output is never the same twice. Adjacent roads count: scoring systems from fraud, search relevance, or content integrity - any domain where ground truth was scarce and the adversary was real. We care about the artifact and the reasoning behind it far more than where you did it.
Who this isn't for. This role is wrong if you optimize for leaderboard scores or trust vendor-reported evals - the work here is deciding what a score even means, not chasing one. It is wrong if your instinct when output looks weak is to massage the prompt rather than redesign the verification, and wrong if you would rather publish a finding than ship a gate. It is wrong if you want tightly-scoped tickets and a lane to stay in; the scope of this seat is the whole verification surface. And it is wrong if you are comfortable shipping what an agent produced without being able to say why it is right, or letting an agent grade its own work. You'll be happiest here if you have high agency, think in corpora, and want your verification to be the reason an entire truth layer can be trusted.
How We Evaluate
We don't run traditional AI engineering interviews. Every stage is demonstrated performance on work-relevant tasks.
Async video screen. Brief and on your own time: about 15 minutes. We want to see how you think, not how you present. Calls with company stakeholders. Short conversations with key members of the team. Conversation with the founder. How you model the system above, where you push back, and whether you can hold the argument live. Paid work trial. One week of paid, real verification work in our real environment - live loops, live corpora, real verdicts. It is paid because your time is worth paying for. We watch how you get grounded, whether you write the spec before the build, how you verify what your agents produce, and whether your self-assessment is honest.
If the work above reads like yours but your resume is unconventional, apply anyway. We hire on the work and the reasoning, not the pedigree.
Compensation & Ownership
Total first-year comp: $400,000 - $500,000 (base + performance-based ownership and profit-share programs). Base: $260,000 - $330,000 - top of market for machine learning engineering.
Beyond base: eligibility for the company's ownership and profit-share programs - grants are performance-based, with terms discussed at the offer stage. 100% premium coverage for you and your family. The token budget is effectively unlimited, steered by ROI and never capped.
This structure is built to mint partners. When the company wins, you win - in real, liquid dollars, every year.
Based in Santa Monica, Los Angeles - in person, five days a week. The rooms are real rooms. Relocation support available for the right builder.
Product.ai is the verified truth layer for shopping: the intelligence that tells you what is actually true about a product, including when not to buy. SimplyCodes is its first proof at scale: the code verification service - codes that actually work, proven by robots running real checkouts - at around $22M in revenue and roughly 60% margins. Founder-owned and profitable since 2009. No outside investors. No board. Fewer than twenty operators outbuilding companies 10x our size.
Strong people find us and keep finding us - they apply over months and years, because the field moves fast and the exact profile we need moves with it.
Why This Role Exists
Everything we sell rests on one question: is the claim true? Machine verdicts answer it - they decide what we publish about a product, a price, a code. Today machines settle a minority of those verdicts and humans settle the rest. The constraint on growth is not more claims; it is whether the machine verdicts stay right as the volume climbs. You own that: the eval science behind every automated verdict - the judges, the thresholds, and the exams that grade them. You work directly with the founder.
The target is the large majority of verdicts settled by machine, holding a high accuracy floor against a held-out, human-graded exam. You co-own the arbitration rulebooks - the rules that settle a contested verdict - with our truth scientist, so the truth an entire company ships through never rests on one person's judgment. It is also the most durable engineering seat in the building: every model generation absorbs another layer of surface skill, but grading a verdict no human can check faster than the machine is an open frontier that compounds for years.
The System You'll Need to Model
- Adversarial robustness, end to end. Merchants, bots, and models all have incentives to game a verdict. A merchant wants its expired code to read as live; a bot wants to look like a shopper; an LLM judge grades its own model family measurably too generously - self-preference bias you measure and correct, never assume away. Your verifier has to be structurally harder to fool than any of them.
- Calibration with decay. A verdict is not a yes or a no; it carries a confidence and a shelf life. A price claim goes stale in days; a materials fact holds for years. The pipeline has to score how sure the machine is, per claim class, and stay honest when it is not sure - calibrated confidence with principled abstention, because a confidently wrong verdict is worse than no verdict at all.
- Golden sets - grading the grader. The hardest problem in the seat is meta-evaluation: measuring a judge whose judgments no human can make faster than the machine. The answer is a held-out, human-graded exam the model cannot see, cannot train on, and cannot game. Label quality, inter-rater agreement, contamination control, refresh cadence - designing that exam is the science.
- Evals as the ship gate. Eval-driven development is not a testing afterthought here; it is how engineering works. Agent loops write the code and the content; your corpora and gates decide whether a loop's output ships or gets rejected. The eval is the spec.
- Cortex - the shared AI brain. You work inside Cortex, the governed AI substrate that runs the company and is the product family we sell; it answers its own questions from 8,600+ documents. Nobody else runs the company on the machine they sell. Your verification is what keeps its answers trustworthy.
- Scale, under model churn. Hundreds of thousands of merchants, millions of claims. An eval method that holds on a hundred cases and falls over at a million is a demo, not a method. And the ground moves under it: each model generation resets what judges can do and what your exams still discriminate, so the harness gets rebuilt while it runs.
If reading that energizes you, keep going. If it feels overwhelming or underspecified, this isn't the right fit.
What You Will Own
- The judge pipeline. The machine verdicts that decide what we claim is true about shopping - the models, the thresholds, and the rulebooks behind every automated verdict. Your number is the share of verdicts machines settle: you take it from a minority to the large majority, holding a high accuracy floor against the human-graded exam.
- Calibration and abstention. How sure the machine is, per claim class, with the decay each class carries. You own the pipeline's honesty: a judge that reports 95% confidence and is right 70% of the time is miscalibrated, and that bug is yours.
- The golden sets. You grow them, you keep them clean - label quality, contamination control, refresh before they saturate - and you give them teeth. When a verdict is disputed, the exam decides, not the loudest engineer in the room.
- The arbitration rulebooks, co-owned. The adjudication rules that settle a contested verdict, authored with our truth scientist - two owners by design.
- Your seat charter. Within your first quarter you co-sign a charter for this seat - one machine-checkable number that proves it is working, and a written split of what you decide freely versus what you bring to the founder to decide. This is the model we run: a real authority split, in writing, not a job description.
You will use the craft you already own - judge design, golden sets, calibration, regression corpora, statistical rigor - and grow into the layer above it: eval-gated agent orchestration, adversarial verification at commerce scale, and adjudication design for judgments no single grader can settle.
Who You Are
You interrogate every green checkmark. A passing eval is a claim, and claims get challenged - you ask what the test could not have caught before you ask what it confirmed. You form working models of complex systems on your own, notice where your model is wrong, and update fast. You write clearly, because clear writing is evidence of clear thought.
You treat agents as leverage you verify, never as an oracle you trust. You can do this job by hand and prove it - hand-grade a judge, hand-build a corpus, hand-verify a claim set - and that mastery is exactly what lets you trust, or reject, the verdict an agent hands back. The expensive thing here is a redo cycle, never the compute.
You have probably built an eval harness another team came to depend on, a regression corpus that caught a real failure before users did, an LLM-judge pipeline where you measured the judge's bias instead of trusting it, or a golden set for a system whose output is never the same twice. Adjacent roads count: scoring systems from fraud, search relevance, or content integrity - any domain where ground truth was scarce and the adversary was real. We care about the artifact and the reasoning behind it far more than where you did it.
Who this isn't for. This role is wrong if you optimize for leaderboard scores or trust vendor-reported evals - the work here is deciding what a score even means, not chasing one. It is wrong if your instinct when output looks weak is to massage the prompt rather than redesign the verification, and wrong if you would rather publish a finding than ship a gate. It is wrong if you want tightly-scoped tickets and a lane to stay in; the scope of this seat is the whole verification surface. And it is wrong if you are comfortable shipping what an agent produced without being able to say why it is right, or letting an agent grade its own work. You'll be happiest here if you have high agency, think in corpora, and want your verification to be the reason an entire truth layer can be trusted.
How We Evaluate
We don't run traditional AI engineering interviews. Every stage is demonstrated performance on work-relevant tasks.
If the work above reads like yours but your resume is unconventional, apply anyway. We hire on the work and the reasoning, not the pedigree.
Compensation & Ownership
Total first-year comp: $400,000 - $500,000 (base + performance-based ownership and profit-share programs). Base: $260,000 - $330,000 - top of market for machine learning engineering.
Beyond base: eligibility for the company's ownership and profit-share programs - grants are performance-based, with terms discussed at the offer stage. 100% premium coverage for you and your family. The token budget is effectively unlimited, steered by ROI and never capped.
This structure is built to mint partners. When the company wins, you win - in real, liquid dollars, every year.
Based in Santa Monica, Los Angeles - in person, five days a week. The rooms are real rooms. Relocation support available for the right builder.
Similar Jobs at Product.ai
Artificial Intelligence • Big Data • Consumer Web • eCommerce
Own and monetize the agent-economy revenue line: price and sell keyed API access, convert free merchant alerts into paid find-and-fix contracts, manage affiliate-network relationships, and build platform partnerships and pilots (ChatGPT apps, App Intents, Gemini, Claude). Deliver a measurable first-quarter charter and signed pilots; run end-to-end commercial motion from packaging and pricing to partner integration.
Top Skills:
APIsApple App IntentsChatgptClaudeCortexGeminiKeyed ApiMcp
Artificial Intelligence • Big Data • Consumer Web • eCommerce
Owner of company-wide operational outcomes: run and verify long-lived AI agents to move measurable metrics (first: hiring candidate-to-decision latency), build operating cadence, run back-office/finance/vendor operations, design an ownership program, and grow a human-workforce platform. You model complex systems, instrument processes, write clear specs, and ship operational products inside the company AI brain (Cortex).
Top Skills:
Ai AgentsCortex
Artificial Intelligence • Big Data • Consumer Web • eCommerce
Own the consumer quality bar and author locked specs across chat, web, extension, mobile, and personalization. Translate founder strategy into falsifiable product outcomes, verify agent-built implementations, define decision-shaped UX with design, and run 1–2 end-to-end measurable outcomes per quarter. Operate as a founding individual contributor, directing agents and elite operators, and co-own an operator contract that defines measurable seat performance.
Top Skills:
Ai AgentsApple App IntentsChatgptClaudeCortexGeminiKnowledge GraphsVerifier Agents
What you need to know about the Colorado Tech Scene
With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.
Key Facts About Colorado Tech
- Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
- Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
- Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
- Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

.png)