The University of Cambridge is recruiting a Research Associate to develop large language model systems that can extract biodiversity information from scientific literature at very large scale. The post sits in the Department of Zoology and combines artificial intelligence, information extraction, ecology, genomics and climate science.
The successful researcher will work on agentic LLM workflows capable of processing more than one million scientific papers and turning unstructured research literature into structured biodiversity evidence. The salary is £37,694–£46,049, funding is available for up to three years, and applications close on 23 August 2026.
Position at a Glance
- University: University of Cambridge
- Department: Zoology
- Reference: PF50528
- Role: Research Associate in Large Language Models for Biodiversity Data Extraction
- Location: Cambridge, United Kingdom
- Work arrangement: On-site; flexible-working requests may be considered
- Salary: £37,694–£46,049
- Funding period: Up to 3 years
- Closing date: 23 August 2026
- Interview timing: Week beginning 31 August 2026
Why Cambridge Is Hiring an LLM Researcher for Biodiversity
Modern biodiversity research depends on several types of evidence at once. Researchers may need species-location records, genomic information, environmental variables, ecological interactions, demographic histories and climate data. A large amount of that information already exists in published papers, but much of it is trapped inside prose, tables, figures and supplementary material rather than organised in a format that modelling systems can use directly.
Cambridge wants to use modern LLM and agentic-AI methods to close that gap. The researcher will design systems that can read scientific literature at scale, extract relevant facts, organise the information and validate it before the data is used in biodiversity forecasting.
The Wider CISGeM Research Programme
The role forms part of work on Climate-Informed Spatial Genomic Models (CISGeMs). These models aim to combine genomic, ecological, climate and environmental information so researchers can reconstruct past population histories and estimate how species and ecosystems may respond to future environmental change.
The LLM component matters because scientific literature contains a huge amount of contextual knowledge that is difficult to capture through existing databases alone. For example, a paper may report a species distribution, a local population decline, a predator-prey relationship, an environmental association or demographic event that could materially improve a forecasting model.
Processing More Than One Million Scientific Papers
The vacancy is unusually explicit about scale. The successful candidate is expected to help develop workflows capable of processing more than one million scientific papers.
This makes the position very different from a small document-classification research project. The system needs to be robust enough to work across scientific fields, writing styles and document structures while maintaining traceability and quality control.
Potential technical challenges include document retrieval, chunking and context management, entity resolution, georeferencing, schema design, tool use, citation tracking, retrieval-augmented generation, uncertainty estimation and automated or semi-automated validation.
What Kind of Information the AI System May Extract
Cambridge identifies several categories of biodiversity information that matter to the programme:
- Species distributions and geographic records.
- Ecological interactions between organisms.
- Population and demographic processes.
- Environmental associations.
- Climate-related biological responses.
- Other evidence relevant to conservation and biodiversity forecasting.
The goal is not merely to summarise papers. The system needs to turn literature into data that can support downstream scientific analysis.
Agentic AI Is Central to the Role
The advert specifically mentions agentic LLM-based systems. In practice, that can involve models that perform multi-step tasks rather than responding to a single prompt. An agent may retrieve a paper, identify relevant sections, call tools, extract candidate facts, validate units or locations, compare evidence across sources and store outputs in a structured representation.
Applicants with experience in tool-using language models, multi-agent workflows, autonomous research pipelines or evaluation of agent behaviour should make that experience highly visible.
Essential Technical Background
Candidates should hold a PhD in a relevant discipline such as:
- Computer Science.
- Machine Learning.
- Artificial Intelligence.
- Bioinformatics.
- Computational Biology.
- A closely related quantitative discipline.
A strong quantitative background is expected. Cambridge also asks for substantial experience in one or more areas including large language models, natural-language processing, information extraction, retrieval-augmented generation and agentic AI systems.
Programming and Machine-Learning Skills
Excellent programming skills and practical experience with modern machine-learning ecosystems are expected. A competitive applicant should be able to demonstrate that they can move beyond a notebook prototype and contribute to shared research infrastructure.
Examples of useful evidence include building reproducible pipelines, maintaining research code, working with cloud or cluster compute, evaluating large models, constructing data-processing systems, contributing to open-source software or supporting other researchers through reusable tools.
Is an Ecology Background Required?
No. Experience in ecology, biodiversity science, geospatial analysis or scientific text mining is advantageous but not essential.
This means a technically strong LLM or NLP researcher can still be a good fit without a formal ecology degree. The key is to show an ability to work with domain scientists and learn the biological context needed to build trustworthy extraction systems.
The Research Group
The successful candidate will join a large interdisciplinary group with more than 20 PhD students and postdoctoral researchers working across ecology, evolution, conservation and genomics.
The LLM researcher will interact with colleagues developing deep-learning methods for population-genomic inference, population geneticists producing genomic datasets and ecological modellers using those tools for biodiversity questions.
This environment makes collaboration an important part of the role. A strong candidate should be comfortable explaining AI methods to non-AI specialists and understanding scientific requirements from researchers in other disciplines.
Weekly Hackathons and Shared Software
The post includes participation in weekly hackathons and collaborative coding sessions. The researcher will also help build shared software infrastructure and open-source tools.
That detail is worth highlighting in an application because it suggests Cambridge is looking for someone who enjoys hands-on development and collective problem solving, not only an individual researcher focused on publications.
Mentoring and AI Leadership
The successful applicant may mentor junior researchers and is expected to share expertise in LLMs and AI across the wider programme.
Applicants who have supervised students, reviewed code, run reading groups, delivered technical workshops or helped colleagues adopt LLM methods should mention that experience.
How to Position Your Research Experience
A generic AI CV may not be enough. The strongest application will connect previous work to the scientific problem Cambridge is trying to solve.
If your background is in NLP, explain how you handled extraction accuracy, entity linking, retrieval, evaluation or difficult scientific text. If your background is in LLM agents, show how you managed multi-step reasoning, tool calls, error recovery or verification. If you come from bioinformatics or computational biology, explain how AI methods interacted with biological data and scientific interpretation.
What to Emphasise in Publications
Highlight work relevant to:
- Large language models.
- Information extraction.
- Retrieval systems.
- Agentic AI.
- Scientific machine learning.
- Bioinformatics.
- Geospatial data.
- Biodiversity or ecology.
- Open-source research software.
If a paper created a system or dataset that other researchers used, explain that impact clearly rather than relying only on publication venue names.
On-Site Requirement
Cambridge states that the role is based entirely on site because of the nature of the work, although flexible-working requests will be considered.
International applicants should also remember that University employees must be eligible to live and work in the United Kingdom. The public vacancy does not make a specific visa-sponsorship promise, so candidates who require sponsorship should check the University’s current immigration guidance and the eligibility of this particular post.
Interview Timing
Interviews are planned for the week beginning 31 August 2026. Because this follows shortly after the application deadline, shortlisted applicants should be prepared to discuss technical work and research plans quickly.
Prepare concise explanations of one or two relevant research projects, including the scientific problem, your technical contribution, evaluation method, limitations and what you would do differently now.
How to Apply
Applications are submitted through the University of Cambridge online recruitment system. Applicants should quote reference PF50528 in the application and any correspondence.
The deadline is 23 August 2026. Because the deadline is immediate, candidates should prioritise a complete and technically focused application rather than spending excessive time on cosmetic formatting.
Frequently Asked Questions
What is the salary for the Cambridge LLM biodiversity role?
The published salary range is £37,694–£46,049.
How long is the position?
Funding is available for up to three years.
Do I need a PhD?
Yes. The role asks for a PhD in computer science, machine learning, AI, bioinformatics, computational biology or a related discipline.
Is ecology experience mandatory?
No. Ecology, biodiversity, geospatial or scientific-text-mining experience is advantageous but not essential.
What is the main AI focus?
The project focuses on agentic LLM systems for extracting, organising and validating biodiversity knowledge from more than one million scientific papers.
