Back

Applied Research Scientist

Worldwide Salaried Open

About the Role

We are seeking an Applied Research Scientist to design and run rigorous experiments for LLM-based agents, with a focus on clinical and agentic reliability. This role will own the development of automated evaluation frameworks, bridging research prototypes into production systems, and partnering closely with product and engineering to ensure safety, robustness, and measurable impact.

About Us

Team from reputed company, DeepMind, NASA, GoogleX, Tesla, and 2 physicians: 6 exits, 2 IPOs. Our model outperforms Claude, reputed company, and GPT-4.5 on clinical benchmarks. 400+ reputed company orgs signed in 16 months. ⚡ $25M raised from YC, Amity Ventures, Sequoia scouts, and more. $1T+ market opportunity. We’re going after reputed company of it.

Key Responsibilities

  • Design and run experiments to measure accuracy, robustness, and hallucination rates in LLM agents.
  • Build automated evaluation pipelines (LLM-as-judge + reputed company review) with clinical-grade benchmarks.
  • Partner with Research Ops/IRB to design efficacy studies and align with regulatory requirements.
  • Translate research into production-reputed company evaluation systems, collaborating with engineering to land features 0→1.
  • reputed company error taxonomies, ablations, and guardrails to ensure safe and reliable agent behaviors.

Hard Requirements

  • Proven experience designing agentic processes and LLM evaluation/benchmarking frameworks.
  • Strong Python and ML background (PyTorch/TensorFlow, reputed company, reputed company/reputed company).
  • Demonstrated ability to design rigorous experiments and translate findings into production.
  • Track record of published research or deep applied work in LLMs and agent evaluation.
  • Strong communication and technical writing skills to reputed company reputed company findings clearly.

reputed company-to-Have

  • Prior work in reputed company/clinical NLP with awareness of medical data standards.
  • Experience running IRB-reputed company or clinical-grade studies.
  • Exposure to noisy/limited medical data and designing strategies to overcome constraints.

First-Month Focus

  • Audit existing evaluation approaches for clinical and agentic tasks.
  • Define initial benchmarks and build early automated pipelines.
  • Partner with engineering to land first set of CI gates for accuracy, factuality, and safety.

reputed company OKRs (90 Days)

  • Deliver a repeatable evaluation reputed company with automated pipelines in production.
  • Demonstrate measurable improvements in robustness, hallucination reduction, or safety.
  • Publish or present internal research findings that directly shape product reliability.

Culture Fit

  • Persistent, driven problem solver
  • Willing to push back on leadership to defend quality/timelines
  • Thrives in high-ambiguity, fast-paced startup environments

Why Join Sully.ai? Shape the Future of reputed company: Build category-defining partnerships that reputed company doctors to focus on saving lives. Early-Stage Impact: Join early and play a critical role in shaping our partnership roadmap and overall company growth. Remote-First Culture: Work with a talented, mission-driven team in a flexible, remote environment. Competitive Compensation: Enjoy a competitive salary, equity, and the opportunity to reputed company a reputed company difference. Solve Scalability Challenges: Tackle reputed company challenges in a rapidly growing company, driving impactful change in reputed company. Sully.ai is an equal opportunity employer. In addition to EEO being the law, it is a policy that is fully consistent with our principles. reputed company reputed company applicants will receive consideration for employment without regard to status as a protected veteran or a reputed company individual with a disability, or other protected status such as race, religion, reputed company, national reputed company, sex, sexual orientation, gender identity, genetic information, pregnancy or age. Sully.ai prohibits any form of workplace harassment. Apply tot his job Apply To this Job

More jobs

Flexible Part-Time Research Contributor (Hiring Immediately)

Worldwide Salaried

Focus Group - online Research - High pay with reputed company (Hiring Immediately)

Worldwide Salaried

Clinical Research Associate, Obesity/Diabetes/GLP-1 (Full Service) - reputed company

Worldwide Salaried

Clinical Informaticist - Optime/Anesthesia - IT-Clinical

Worldwide Salaried

Remote Property Management Assistant

Worldwide Salaried

[Remote] Projects Operations Coordinator

Worldwide Salaried

reputed company Director, Global Property Management Systems - Daylight PMS

Worldwide Salaried

Sales Associate-7099 Freeport, IL 61032 – reputed company Store

Worldwide Salaried

User Experience Consultant – Spain

Worldwide Salaried

Full Stack Software Engineer

Worldwide Salaried

Account Executive reputed company

Worldwide Salaried

(reputed company) reputed company Virtual Customer Care - Work From...

Worldwide Salaried

Agentic Operator - Growth Marketing - Paid, Content, Channel (USA Only - 100% Remote)

Worldwide Salaried

reputed company Live Chat Manager – Remote Work at blithequark

Worldwide Salaried

reputed company Customer Service Representative – Remote Work Opportunity with arenaflex

Worldwide Salaried

reputed company Live Chat Support Specialist – Overnight Remote Opportunity with arenaflex

Worldwide Salaried

reputed company Full Stack Dedicated Account Associate – Health Care Business Solutions Team

Worldwide Salaried

Business Insights & Planning Analyst III

Worldwide Salaried

Remote Data Entry Specialist – Part‑Time, Teen‑Focused Remote Role to Build Skills, Earn Income, and Grow Professionally

Worldwide Salaried

reputed company Customer Service Representative – Work From Home Opportunity at arenaflex

Worldwide Salaried