Senior Applied Scientist
The OCI AI Evaluation team builds the evidence behind model-selection, product-readiness, and launch decisions. We evaluate frontier foundation models and AI systems across capabilities such as reasoning, coding and agentic coding, retrieval-augmented generation, AI agents, NL2SQL, multimodal understanding, multilingual performance, and responsible AI.
As a Senior Applied Scientist on the team, you will independently own complex evaluation work from problem definition through final recommendation. You will translate ambiguous product and customer questions into measurable hypotheses, select or create appropriate benchmarks, design experiments, build evaluation pipelines, validate data and metrics, analyze failure modes, and communicate conclusions to science, engineering, product, and leadership stakeholders.
This is hands-on applied science. You will write high-quality code, work with large and imperfect datasets, develop and calibrate automated evaluators, and turn one-off analyses into reproducible evaluation protocols and reusable infrastructure. You will examine more than aggregate benchmark scores, considering factors such as statistical validity, data provenance, contamination, robustness, cost, latency, reliability, safety, and operational constraints.
The work sits at the point where research results become product decisions. Success requires scientific rigor, strong engineering judgment, clear writing, and the ability to make progress when requirements, model access, data, or infrastructure are still evolving. You will collaborate closely with other scientists, software engineers, product teams, data and human-annotation teams, and external partners to deliver evaluation results that are technically defensible and useful in practice.
Qualifications
Minimum Job QualificationsEducation and/or Experience:
8 years of experience in data science, machine learning, artificial intelligence, natural language processing, speech recognition, statistical modeling, data mining, or related field
OR
Bachelor's Degree in Mathematics or Statistics, Computer Science, Data Science, Physics, or related field AND 4 years of experience in data science, machine learning, artificial intelligence, natural language processing, speech recognition, statistical modeling, data mining, or related field
OR
Master's Degree in Mathematics or Statistics, Computer Science, Data Science, Physics, or related field AND 2 year of experience in data science, machine learning, artificial intelligence, natural language processing, speech recognition, statistical modeling, data mining, or related field.
Job Skills:
Same skills as prior level.
Programming Language:
2 years of experience in applicable programming language.
Preferred Job Qualifications
Education and/or Experience:
9 years of experience in data science, machine learning, artificial intelligence, natural language processing, speech recognition, statistical modeling, data mining, or related field
OR
Bachelor's Degree in Mathematics or Statistics, Computer Science, Data Science, Physics, or related field AND 5 years of experience in data science, machine learning, artificial intelligence, natural language processing, speech recognition, statistical modeling, data mining, or related field
OR
Master's Degree in Mathematics or Statistics, Computer Science, Data Science, Physics, or related field AND 3 years of experience in data science, machine learning, artificial intelligence, natural language processing, speech recognition, statistical modeling, data mining, or related field
OR
Doctorate in Mathematics or Statistics, Computer Science, Data Science, Physics, or related field AND 1 year of experience in data science, machine learning, artificial intelligence, natural language processing, speech recognition, statistical modeling, data mining, or related field.
Responsibilities
Responsibilities
- Independently own end-to-end evaluations of foundation models, AI agents, and enterprise AI systems, from initial question and experiment design through analysis, reporting, and stakeholder review.
- Translate customer, product, and business needs into testable hypotheses, evaluation criteria, datasets, metrics, baselines, and acceptance thresholds.
- Design, implement, and maintain benchmarks and evaluation methods for areas such as reasoning, coding, agentic workflows, RAG, NL2SQL, multimodal systems, multilingual performance, and responsible AI.
- Write high-quality Python and production-oriented evaluation code; build reproducible pipelines, test suites, automated checks, and integrations with shared evaluation platforms.
- Evaluate model and system behavior across quality, cost, latency, reliability, safety, robustness, and domain fit rather than relying only on aggregate scores.
- Conduct statistical analysis, error analysis, ablations, and qualitative failure-mode investigations to explain model behavior and identify meaningful differences between systems.
- Develop and validate automated evaluators, including LLM-as-a-judge methods; calibrate them against human judgments and quantify their reliability, bias, and limitations.
- Design human-evaluation and annotation workflows, including rubrics, gold datasets, sampling plans, quality controls, and vendor or Human-in-the-Loop validation.
- Assess dataset quality, provenance, representativeness, contamination risk, licensing constraints, privacy, and other factors that could invalidate an evaluation or limit use of its results.
- Produce concise, decision-ready reports that make methods, assumptions, limitations, tradeoffs, and recommendations explicit for technical and non-technical audiences.
- Partner with science, engineering, product, data, and operations teams to define requirements, resolve blockers, manage dependencies, and establish clear handoffs and ownership.
- Turn successful evaluation work into reusable protocols, documented workflows, and shared infrastructure that improve the speed and consistency of future evaluations.
- Stay current with research in machine learning, generative AI, agent evaluation, and measurement methodology; prototype promising approaches and contribute to science plans, papers, patents, or technical reports where appropriate.
- Publish original research in top-tier peer-reviewed conferences and journals, and translate relevant evaluation advances into reusable methods, technical reports, or production capabilities for OCI.
- Review technical work, share expertise, and mentor junior scientists or engineers in experimental design, evaluation methodology, coding, and interpretation of results.
- Own delivery quality and timelines, communicate risks early, and maintain clear, auditable documentation of experimental configurations, data versions, results, and decisions.
Minimum qualifications
- PhD in Computer Science, Machine Learning, Artificial Intelligence, Statistics, Mathematics, or a related quantitative field; or a Master’s or Bachelor’s degree with equivalent relevant industry experience.
- Experience designing and executing machine learning experiments, including defining hypotheses, selecting datasets and metrics, establishing baselines, and interpreting results.
- Strong knowledge of modern machine learning, deep learning, natural language processing, and generative AI methods.
- Hands-on experience evaluating large language models, foundation models, AI agents, or other machine learning systems.
- Proficiency in Python and experience writing reliable, maintainable code for experiments, data processing, or machine learning systems.
- Experience working with machine learning frameworks and data-science libraries such as PyTorch, TensorFlow, Hugging Face, NumPy, pandas, or equivalent tools.
- Demonstrated ability to analyze complex datasets, identify data-quality problems, and perform statistical, error, and failure-mode analysis.
- Experience translating technical findings into clear written recommendations for science, engineering, product, or business stakeholders.
- Ability to work independently on complex assignments while collaborating effectively with scientists, engineers, product managers, and other cross-functional partners.
- Strong written and verbal communication skills.
Preferred qualifications
- Experience building evaluation frameworks, benchmarks, test suites, or observability systems for generative AI models and applications.
- Experience evaluating one or more of the following: AI agents, agentic coding systems, RAG, NL2SQL, multimodal models, multilingual models, reasoning systems, or responsible AI capabilities.
- Experience developing or calibrating LLM-as-a-judge methods against human judgments.
- Experience designing human-evaluation programs, including annotation rubrics, sampling plans, gold datasets, inter-annotator agreement analysis, and quality-control processes.
- Experience converting research prototypes or one-off experiments into reusable, production-quality tools and workflows.
- Familiarity with model-evaluation concerns such as benchmark contamination, data provenance, robustness, statistical significance, safety, latency, cost, and reproducibility.
- Experience deploying or integrating machine learning components in cloud or production environments.
- Experience working with distributed computing, large-scale datasets, model serving, or parallel evaluation workloads.
- Publications in top-tier machine learning, natural language processing, data mining, or artificial intelligence conferences or journals.
- Demonstrated ability to identify new research questions, develop novel evaluation methods, and translate research advances into practical capabilities.
- Experience mentoring junior scientists or engineers and providing technical or scientific review.
- Familiarity with enterprise AI requirements, including privacy, security, licensing, compliance, reliability, and auditability.
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.