Axle Informatics Logo

Axle Informatics

Senior Data Scientist, AI Retrieval Systems

Posted 29 Days Ago
Remote
Hiring Remotely in USA
130K-150K Annually
Senior level
Remote
Hiring Remotely in USA
130K-150K Annually
Senior level
Build production AI retrieval and knowledge systems for rare disease research at NIH. Model biomedical ontologies in PostgreSQL, develop semantic search, retrieval-augmented services, ranking, evaluation, and observability, and deliver user interfaces with Next.js, React, and TypeScript. Deploy and operate services in on-premises and HPC Kubernetes environments, including GPU inference. Collaborate with NIH staff, clinicians, and researchers, and contribute to scientific publications.
The summary above was generated by AI

Axle is a bioscience and information technology company that offers advancements in translational research, biomedical informatics, and data science applications to research centers and healthcare organizations nationally and abroad. With experts in biomedical science, software engineering, and program management, we focus on developing and applying research tools and techniques to empower decision-making and accelerate research discoveries. We work with some of the top research organizations and facilities in the country including multiple institutes at the National Institutes of Health (NIH).


Benefits We Offer:

  • 100% Medical, Dental & Vision Coverage for Employees
  • Paid Time Off and Paid Holidays
  • 401K match up to 5%
  • Educational Benefits for Career Growth
  • Employee Referral Bonus
  • Flexible Spending Accounts:
    • Healthcare (FSA)
    • Parking Reimbursement Account (PRK)
    • Dependent Care Assistant Program (DCAP)
    • Transportation Reimbursement Account (TRN)

Axle is seeking a Senior Data Scientist, AI Retrieval Systems to join our vibrant team supporting rare disease research at the National Institutes of Health (NIH). This is a Remote position within the United States. 

 

Position Summary:

Roughly 25 to 30 million people in the United States live with a rare disease. There are somewhere between 7,000 and 10,000 distinct rare conditions, and the large majority have no FDA-approved treatment.

Research on these conditions keeps running into the same obstacles. Published evidence for any one disease is thin and scattered across sources. The same clinical finding gets written down a dozen different ways depending on who recorded it. And the people with the most at stake, patients and their families, are usually the least equipped to read the specialist literature written about their own condition.

Large language models are well suited to this class of problem, and the research programs we support are investing in applying them carefully. In this role you will build the retrieval and knowledge layer that those AI systems stand on. That means the disease and phenotype vocabularies that give a model something precise to reason over, the semantic search that finds the right concept behind an imprecise human phrase, and the ranking that decides what a user sees first. Ontologies serve as internal scaffolding throughout. Users should never have to see one or learn what it is.

This is a senior individual contributor position with unusual range. You will own the data layer, the retrieval services built on top of it, the interfaces where results become visible, and the path onto the computing infrastructure that runs it all. You will work directly with NIH program staff, clinical geneticists, and rare disease information specialists.

 

Core Responsibilities:

  • Model biomedical knowledge for rare disease research. Ingest disease and phenotype ontologies and controlled vocabularies into PostgreSQL with a maintainable release and refresh path, reconcile identifiers across sources, and work through term hierarchies to determine what is clinically relevant for a given condition.
  • Build retrieval-augmented services that ground everyday language in clinical concepts. Embed term labels, definitions, and synonyms, retrieve candidates, and have a model disambiguate against context before any value is committed.
  • Treat retrieval as a database problem. Tune keyword and vector search over large biomedical corpora, and be ready to defend the recall and latency trade-offs you choose.
  • Build the ranking and relevance layers that decide what surfaces first, including domain-aware weighting and graceful degradation when a condition falls outside curated coverage.
  • Deliver the interfaces where this work becomes visible to users, in Next.js, React, and TypeScript. This covers question and confirmation flows, result presentation, and live status for long-running pipelines.
  • Deploy continuously onto NIH on-premises and high-performance computing Kubernetes environments. Helm charts, StatefulSets, secrets, ingress, GPU scheduling for self-hosted inference, and scheduled jobs are all in scope, and you will partner with the operations teams that run those environments instead of standing up parallel cloud infrastructure.
  • Build the evaluation that tells us whether retrieval and concept mapping are good enough to rely on, and keep it running as a regression suite instead of a one-time measurement.
  • Log what the system does and why. Request identifiers, latency, errors, and which concept the system selected all need to be captured, so that staff can review an AI-assisted result instead of taking it on faith.
  • Work out what researchers, clinicians, and patient communities need, and turn it into data models, retrieval behavior, and interface design.
  • Write the work up. You will contribute to manuscripts, conference abstracts, and posters with NIH investigators, and you will be credited as an author on work you helped produce.

 

Required Qualifications:

  • Bachelor’s degree in Data Science, Computer Science, Bioinformatics, Biomedical Informatics, or a related field. An advanced degree is preferred. We will consider equivalent professional experience in place of a degree.
  • At least 5 years building and operating production software or data systems. At least 2 of those years should involve shipping LLM-powered applications (agents, retrieval, or evaluation) that people depend on. We weigh depth in retrieval and applied LLM engineering more heavily than total years.
  • Experience building retrieval systems end to end, covering indexing, query construction, and measuring retrieval quality against real data.
  • Experience evaluating systems that have no single right answer, using golden sets, offline regression suites, or metrics such as Recall@K and MRR to decide whether a change was an improvement.
  • Experience with structured output and tool or function calling, meaning you have constrained a model to a typed schema and validated what came back.
  • Ability to own a service end to end, from schema design through deployment and operation.
  • Ability to obtain and maintain a Public Trust Security clearance.

 

Technical Skills:

  • Python, with FastAPI, Pydantic, and pytest.
  • PostgreSQL at depth, covering vector search (pgvector or equivalent), full-text search, embedding pipelines, indexing, and query tuning.
  • LLM application engineering: provider APIs and gateways, prompt and context design, structured generation, and tool use.
  • Data ingestion and transformation pipelines with a repeatable refresh path.
  • Containers and Kubernetes, enough to ship, debug, and operate a service on infrastructure you do not administer.
  • Working comfort in Next.js, React, and TypeScript.
  • Git-based collaboration and CI/CD in a shared codebase.

Preferred Skills:

  • Biomedical ontologies and controlled vocabularies, including MONDO, HPO, UMLS, MeSH, and other OBO Foundry resources, along with comfort working through term hierarchies, synonyms, and cross references.
  • Grounding model output in a domain terminology through embedding-based retrieval plus model disambiguation, such as entity linking, concept normalization, or ontology alignment.
  • Helm, and deployment to on-premises or HPC Kubernetes environments.
  • Serving open-weight models in production with Ollama or vLLM behind a gateway such as LiteLLM, and work with domain embedding models such as MedCPT.
  • Background in rare disease, clinical genetics, or translational research.
  • Contributions to an open biomedical resource, standard, or consortium, such as OBO Foundry ontologies or GA4GH.'
  • Published or presented work that explains your engineering to people who did not build it. Peer-reviewed papers, conference talks, preprints, technical blog posts, and public open source contributions all count.
  • Prior or current NIH experience.

We are looking for an engineer first. If you have shipped retrieval systems that people depend on and have never opened an ontology file, we want to hear from you. Rare disease and ontology background is useful but not required, and we expect to teach the domain to whoever we hire. Candidates who meet the required qualifications and none of the preferred ones are encouraged to apply.


Disclaimer: The above description is meant to illustrate the general nature of work and level of effort being performed by individuals assigned to this position or job description. This is not restricted as a complete list of all skills, responsibilities, duties, and/or assignments required. Individuals may be required to perform duties outside of their position, job description or responsibilities as needed.

The diversity of Axle’s employees is a tremendous asset. We are firmly committed to providing equal opportunity in all aspects of employment and will not tolerate any illegal discrimination or harassment based on age, race, gender, religion, national origin, disability, marital status, covered veteran status, sexual orientation, status with respect to public assistance, and other characteristics protected under state, federal, or local law and to deter those who aid, abet, or induce discrimination or coerce others to discriminate.

Accessibility: If you need an accommodation as part of the employment process please contact: [email protected]

This role has a market-competitive salary with an anticipated base compensation range listed below. Actual salaries will vary depending on a candidate’s experience, qualifications, skills, and location.


Salary Range
$130,000—$150,000 USD

Similar Jobs

25 Minutes Ago
In-Office or Remote
140K-185K Annually
Senior level
140K-185K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Leads Circle’s third-party security program by conducting complex partner and vendor assessments, evaluating technical controls, identifying risks, and communicating recommendations. Builds AI-enabled tooling, automation, monitoring, scoring, and intelligence-driven workflows to modernize security reviews and support continuous risk management. Partners with Security Engineering, Legal, Procurement, and business teams to embed security requirements into third-party onboarding and oversight.
Top Skills: AIApi SecurityBlockchainCloud SecurityCryptographyDoraGenerative AiIdentity And Access ManagementIso 27001Network ArchitectureSecurity AutomationSmart ContractsSoc 2
An Hour Ago
Remote or Hybrid
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Lead functional safety consulting for vehicle, autonomous systems, industrial automation, battery, and energy storage programs. Develop safety concepts, requirements, safety cases, hazard analyses, and hardware and software safety work products. Perform FMEA, FMEDA, FTA, STPA, and DFA assessments while ensuring compliance with functional safety standards. Lead workshops, training, client reviews, technical assessments, and confirmation measures. Support business development through statements of work, estimates, proposals, and industry representation. Occasional travel to client sites may be required.
Top Skills: DfaFmeaFmedaFtaHaraIec 61508Iec/En 62061Iso 13849Iso 21448Iso 26262Iso 61511StpaUl 4600
An Hour Ago
Remote or Hybrid
110K-150K Annually
Senior level
110K-150K Annually
Senior level
Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Leads automotive and embedded-systems cybersecurity consulting engagements, developing ISO/SAE 21434 and ISO 24089 processes, templates, risk analyses, cybersecurity concepts, verification documents, and assessment reports. Guides clients in secure software design, coding, testing, and compliance with UN Regulations 155 and 156. Responsibilities also include client program leadership, proposals and business development, standards participation, technical guidance, and occasional travel to client sites.
Top Skills: Aspice For CybersecurityCybersecurity Management Systems (Csms)Cybersecurity Resilience Act (Cra)Dynamic Application Security Testing (Dast)Embedded SystemsFuzz TestingIso 21448 (Sotif)Iso 24089Iso 26262Iso 27001Iso/Sae 21434Network SegmentationNist Cybersecurity FrameworkOwaspPenetration TestingSecure BootSoftware Update Management Systems (Sums)Static Application Security Testing (Sast)Un Regulations No. 155 And 156Vulnerability Scanning

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account