Applied AI Evaluation Scientist

Jump App · Remote · FullTime

Apply on company site

Applied AI Evaluation Scientist

Location: Remote (U.S.)

Team: AIML Quality — reporting into Engineering leadership

Level: Senior (IC)

About Jump

Jump's mission is to empower financial advisors, firms, and clients to thrive in the age of AI. We automate meeting prep, note-taking, compliance documentation, CRM updates, client recaps, and follow-up tasks — allowing advisors to process meetings in minutes, not hours. Since launching in January 2024, Jump has grown to 30,000+ users at firms ranging from solo practitioners to enterprise RIAs and independent broker-dealers, including partnerships with LPL Financial, Sanctuary Wealth, Osaic, and others.

Jump is a Series A company, having raised $30M in venture capital from Battery Ventures (lead), Citi Ventures, Sorenson Capital, and Pelion Venture Partners. Our team of 100+ includes leaders from Google, Stripe, JP Morgan, Snowflake, Fidelity, BILL, Apple, Harvard, Stanford, and other top companies and schools.

Our team values: Velocity · World Class · Direct and Kind with No Drama

About the Role

We're looking for an Applied AI Evaluation Scientist — someone who sits at the intersection of data science, information retrieval, machine learning, and product thinking. This person will own the quality and trustworthiness of our AI/ML systems by designing, building, and running rigorous evaluation frameworks. The primary focus will be on our Agentic Retrieval-Augmented Generation (RAG) pipelines — optimizing how we chunk, embed, retrieve, rank, and generate — but the role extends to evaluating other AI/ML systems across the company.

The ideal candidate has the judgment to know what's worth evaluating, what isn't, and the statistical grounding to make sure the evaluations they do run are sound, realistic, and actionable. Balancing resource capacity and velocity is key--knowing what to measure and how to measure it to drive improvements for our customers is paramount.

You will work closely with Product and Engineering. Your code doesn't need to be production-hardened, but it must achieve intended outcomes — think research-quality Python, clear notebooks, and reproducible experiments, not bulletproof microservices.

What You'll Do

Agentic RAG Pipeline Evaluation & Optimization (Primary Focus)

Broader AI/ML Evaluation

Collaboration & Data Review

What We're Looking For

Must-Have

Nice-to-Have

What This Role Is Not

Why This Role Matters

Every team building with LLMs hits the same wall: "Is this actually working?" Most teams guess. We want to know — with evidence, grounded in real data, measured with methods we can defend. This role is the difference between shipping AI features on vibes and shipping them with confidence.

Qualifications

Why Jump

Jump is an equal opportunity employer. We value diversity and are committed to creating an inclusive environment for all employees.

Job alert

Get new jobs by email

Save this search and get relevant new jobs when they appear.