F. Zorrilla

About

I have worked in metagenomics and metabolic modelling for 7+ years, turning genomes and metagenomes into predictions of what each microbe in a community consumes and produces, together with wet-lab partners in academia and industry. At ETH Zürich I co-led a multi-institution NCCR Microbiomes flagship project and mentored three PhD students, and I have taught these methods in five courses.

Looking for my next role

Roles
Senior or Lead positions in computational biology, bioinformatics or scientific data
Sectors
Biotech, pharma, food or agtech
Location
Switzerland, Zürich area preferred
Start
From December 2026, flexible on the exact date
Permit
Italian citizen, Swiss B permit
Languages
English and Spanish (native level), Italian (B1), French (A1), German (A1)
Skills
Metagenomics and microbial genome analysis · genome-scale metabolic modelling · reproducible pipelines (Snakemake, Conda, Slurm) · R and Python · protein language models and structure search (ProstT5, Foldseek)

How I could help an R&D team

  • Predict from genomes which nutrients the strains in a collection cannot make and which partners could supply them, and check those predictions against growth data, including where they fail (Nature Microbiology 2026).
  • Model how the strains in a starter culture or consortium feed each other, alongside fermentation and omics data (Nature Communications 2023, with Chr. Hansen).
  • Build a draft genome-scale metabolic model of a production organism from its genome (Hansenula protocol, 2022).
  • Annotate genes that sequence similarity misses, using protein language models and structure search (ProstT5, Foldseek), and feed the results into metabolic models (ETH Zürich, with EPFL).
  • Build reproducible analysis pipelines that other groups install and run on their own data (metaGEM).
  • Turn a broad research goal into a plan with clear modules, bring the right people in to lead them, and keep them on track (ETH Zürich).

Work

13 publications

8 journal papers, 1 book chapter, 3 preprints and my PhD thesis. Open a card to see the problem, the approach and my part in it. Code, models and data from metaGEM, the soil and cheese papers and my thesis are public on GitHub and Zenodo. Citations from Google Scholar, Sept 2026.

Selected papers

Problem

Community metabolic models were usually built by matching 16S profiles or recovered genomes to reference genomes and reusing the references' models, which misses the strains actually present in a sample. metaGEM builds sample-specific models directly from the genomes assembled from each metagenome.

My contribution

First author and lead developer. I co-conceived and developed the project, ran all the analyses and co-wrote the paper. I wrote the documentation and maintain the Bioconda package. Its users, releases and setup are in the metaGEM section.

Links

Problem

Ecological theory predicts that many soil bacteria cannot make some essential nutrients and rely on neighbours for them, which would also help explain why most soil bacteria will not grow alone in the lab. This had not been tested systematically in real soil communities. For agtech, a strain that relies on neighbours for nutrients may not establish on its own.

Approach

The Kost lab isolated 6,931 bacteria from 27 soil communities in Germany and tested which could grow without added nutrients. We sequenced 62 strains and built metabolic models for them to see whether their genomes predict these dependencies. The lab then grew auxotrophs with strains from the same soil, in pairs and small groups, and the models predicted which neighbours could supply the missing nutrients.

My contribution

Co-first author, shared with Ghada Yousif. I co-designed the genome sequencing and modelling and carried out most of the genomic analysis and metabolic modelling: sequence analysis with metaGEM, genome-scale models of the 62 sequenced strains and their comparison with the Kost lab's growth tests, the insertion-sequence and gene-loss analysis, pangenomics, the phylogenetic regression and the community simulations. I also co-wrote the first draft. The auxotrophy screen was the Kost lab's experimental work. The models missed most auxotrophies, and the genome analysis linked many missed cases to insertion sequences and gene loss in pathway genes, which became a result of the paper. The code repository maps every figure and table to the code and data that produce it.

What I used

metaGEM, CarveMe, SMETANA, ReFramed, Pangenomics, R (phyloglm), Python

Press coverage

Links

Problem

Cheddar starter cultures mix Streptococcus thermophilus with several Lactococcus strains. How these strains interact during ripening, and how that shapes flavour, was largely unknown beyond pairwise tests in simplified lab media.

Approach

The team, led by first author Chrats Melkonian (Chr. Hansen and Wageningen University), made Cheddar with the full commercial starter culture and with single strains left out, and followed it through a year of ripening; shorter milk fermentations left out single Lactococcus strains. They combined genome sequencing, genome-scale metabolic models, metatranscriptomics and metabolite measurements. S. thermophilus breaks down milk protein, which relieves nitrogen limitation for Lactococcus, and a Lactococcus cremoris strain keeps the off-flavour compounds diacetyl and acetoin in check.

My contribution

Second author. I carried out most of the published metabolic modelling, together with first author Chrats Melkonian: genome-scale metabolic models of the starter strains built with metaGEM's modelling module (CarveMe, SMETANA), and the community simulations. This built on earlier modelling by Daniel Machado. I also helped write the paper. The cheese-making experiments were run at Chr. Hansen.

What I used

metaGEM, CarveMe, SMETANA, ReFramed, COBRApy, Python

Press coverage

Links

Problem

Recurrent C. difficile infection responds to faecal transplant, but the active ingredients of "colonisation resistance" are unclear. Which metabolic interactions actually keep the pathogen out?

Approach

Joy Scaria's lab (Oklahoma State and South Dakota State universities) challenged a 14-strain gut community, and communities grown from stool samples, with antibiotics and then C. difficile in a continuous-flow bioreactor; the Patil lab in Cambridge carried out the modelling. Population models and genome-scale metabolic models pointed to interactions between E. coli and Bacteroides/Phocaeicola, and metabolomics pointed to fructooligosaccharide use, vitamin B3 synthesis and competition for the amino acids C. difficile ferments. Patient metagenomes supported these signals.

My contribution

Third author. With Arianna Basile I annotated the genomes and built and simulated genome-scale metabolic models of the community members and C. difficile (CarveMe, SMETANA). I analysed the metabolomics data comparing communities that suppressed C. difficile with those that did not, checked the genomes for the genes behind those differences, and helped analyse the patient metagenomes.

What I used

CarveMe, KEMET, SMETANA, ReFramed, R

Links

Problem

Building and curating a genome-scale metabolic model is slow and tedious, and Hansenula polymorpha (Ogataea polymorpha), an industrially relevant methylotrophic yeast, had no published high-quality model.

Approach

A step-by-step protocol chapter that builds the model from the models of related yeasts with the RAVEN toolbox (MATLAB), and releases the result as hanpo-GEM.

My contribution

First author, with Eduard Kerkhoven (Chalmers). We built the draft model together as a side project during my MSc, and I wrote the protocol chapter; I published the model's v1.0 release on GitHub in 2024.

Tools

RAVEN, MATLAB, SBML, Homology-based reconstruction

Links

Problem

Plastic pollution is growing faster than the list of enzymes known to break plastic down, and finding new ones in the lab is slow. The question was whether global metagenomic data could point to candidates.

Approach

The team (first author Jan Zrimec, Chalmers) built profile hidden Markov models from 95 experimentally verified plastic-degrading enzymes and searched ocean and soil metagenomes from 236 sites, finding over 30,000 candidate enzymes for 10 plastic types. Enzyme counts correlated with measured ocean plastic pollution and, for soils, with country-level data on poorly managed plastic waste.

My contribution

Fourth author. I generated all the metagenome assemblies and genome bins (MAGs) for the study with metaGEM; Jan Zrimec then ran the enzyme searches and all the statistics.

What I used

metaGEM, MEGAHIT, CONCOCT, MetaBAT2, MaxBin2, metaWRAP

Press coverage

Links

Other research (6 papers)
  • Drug tolerance via cross-feeding

    Nat. Microbiol. 2022 · Third author of 19 · 144 citations · Paper (Nat. Microbiol.)

    I carried out the analysis that opens the paper: how common amino-acid auxotrophs are across more than 12,000 Earth Microbiome Project communities, and whether auxotrophs grow better under drugs.

  • Enterobacteriaceae gut dynamics

    Nat. Microbiol. 2025 · Third author · 97 citations · Paper (Nat. Microbiol.)

    I gave the authors a session on flux balance analysis and genome-scale metabolic models, then reviewed the finished paper and helped interpret the metabolic scores.

  • FROG: GEM reproducibility

    bioRxiv 2024 · Co-author · 15 citations · bioRxiv preprint

    As a community contribution, I submitted a genome-scale model from one of our lab's published studies to BioModels as a test case.

  • CarveMe-GutMicrobes

    bioRxiv 2026 · Co-author · bioRxiv preprint

    Automated metabolic model reconstruction for gut bacteria and archaea, from the Patil lab. I contributed to the original idea for the project, which Arianna Basile carried out, and helped with visualisation.

  • iGEM interlab measurement studies

    Commun. Biol. 2020 and PLOS ONE 2021 · Consortium author · 381 and 16 citations · Paper (Commun. Biol.) · Paper (PLOS ONE)

    As a master's student on the Chalmers iGEM 2018 team, I was one of three members in charge of our lab's interlab contribution: growing the test strains, calibrating OD and fluorescence on the plate reader, and counting colonies.

  • PhD thesis

    University of Cambridge 2024 · Full thesis (open access) · GitHub

    Omics-driven and constraint-based modelling of microbial community metabolism.

Browse by collaborator

13 key co-authors and the papers we share. Tap or hover a dot for details; tap a paper to open its summary.

Co-author Paper Industry collab

metaGEM

Open-source software

Builds genome-scale metabolic models directly from metagenomes.

Other groups install metaGEM from Bioconda and run it on their own data. I have supported them since 2021, with six releases and replies to 88 of the 90 GitHub issues other users opened.

The pipeline runs quality control (fastp), assembly (MEGAHIT), binning (CONCOCT, MaxBin2, MetaBAT2), bin refinement and reassembly (metaWRAP), taxonomy (GTDB-Tk), metabolic reconstruction (CarveMe), model QC (MEMOTE) and community simulation (SMETANA), on a workstation or a Slurm cluster.

Install

git clone https://github.com/franciscozorrilla/metaGEM.git
cd metaGEM/workflow
mamba env create -n metagem -f envs/metaGEM_env.yml

Linux (clusters or workstations). metaWRAP, the modelling tools (CarveMe, MEMOTE, SMETANA) and the reference databases need a few more steps: see the setup guide.

Worked example: metaGEM on Unseen Bio gut samples

Applications in the wild

Published studies in which other groups used metaGEM, or followed its modelling steps, on their own data.

  • 2026 · Nature Communications

    Continental-scale soil carbon decomposition

    Song et al. (Pacific Northwest National Laboratory, with co-authors at the University of Arizona, the DOE Joint Genome Institute and Eawag) used metaGEM for quality control, assembly and initial binning of metagenomes from 47 US soil samples (344 of their 828 genomes), and linked the genomes to soil organic-matter chemistry.

    Carbon cycling · climate

  • 2026 · Cell Host & Microbe

    Soil protists and bacterial cooperation

    Liu et al. (Nanjing Agricultural University, with Wageningen University) extended metaGEM with the SemiBin binning tool to build metabolic models from tomato rhizosphere and soil microcosm metagenomes. Simulations with these models were one of several checks behind their main finding: predation by soil protists shifts bacteria from competing toward cooperating.

    Soil ecology · agtech

  • 2025 · Environmental Microbiology Reports

    Subsurface H₂ storage microbiology

    Tinker et al. (US national laboratories NETL, PNNL and Sandia) ran metaGEM with default settings to recover nine draft genomes from an enriched water sample from a deep saline aquifer in southern Illinois, to check whether its microbes could consume hydrogen stored there.

    Energy · underground hydrogen storage

  • 2024 · Communications Biology

    Microbiome for industrial phenolic wastewater

    Zhao et al. (Tianjin Institute of Industrial Biotechnology, CAS) used metaGEM to recover 164 draft genomes from a microbial community they scaled up from shake flasks to three industrial sites treating phenolic resin wastewater, sampling it at six stages along the way.

    Industrial wastewater · biotech

Show 2 more
  • 2022 · F1000Research

    NEON soil metagenomes pipeline

    Werbin et al. (Boston University) moved the main workflow of their public tutorial for the National Ecological Observatory Network's soil metagenomes to metaGEM, citing its support for computing clusters, and thanked the metaGEM developers for help troubleshooting.

    Soil · large-scale ecology

  • 2022 · Environmental Microbiome

    Movile Cave: chemoautotrophic ecosystem

    Chiciudean et al. (Babeș-Bolyai University, Romania) followed metaGEM's modelling steps (CarveMe, MEMOTE, SMETANA) on their own genomes from a sulfidic cave, to map competition and cooperation in an ecosystem that runs without sunlight. Their paper acknowledges my help with the simulations and statistics.

    Extreme environments · ecology

Selected from published studies that use metaGEM; the 150+ works that cite it are listed on Google Scholar.

Curriculum vitae

Postdoc, ETH Zürich (2024 to 2026) · PhD, University of Cambridge · EMBL Heidelberg

Download CV (PDF)

Experience

Postdoctoral Researcher · Co-lead, NCCR Microbiomes Work Package 5 Flagship Project

Sunagawa Lab, ETH Zürich · October 2024 to September 2026

  • Planning: designed the 4-year computational roadmap and its modules for the flagship (2024 to 2028; ETH Zürich, EPFL, UZH, UNIL and CHUV) and pitched it to the PIs.
  • Leadership: coordinated the flagship with a PhD co-lead from EPFL on behalf of the work-package leaders. I co-organised its launch workshop (Bern, March 2025), recruited the co-leads of the four modules, ran about 20 coordination meetings and the progress reports, and co-led two modules until September 2026, with 50+ working sessions. I am co-first author on both module manuscripts (in preparation).
  • Module 4, structure-based metabolic models (with EPFL): the module adds protein-structure predictions to genome-scale metabolic model reconstruction and benchmarks the result against existing methods and experimental data. I chose ProstT5 and Foldseek over AlphaFold so annotation scales to microbiome gene sets, scripted and deployed GPU-enabled annotation on ETH's Euler cluster (Snakemake, Slurm) with custom enzyme and transporter databases, and designed the comparison against sequence search (MMseqs2, DIAMOND) and experimental growth phenotypes.
  • Module 3, machine learning for community assembly (with UZH): the module trains a neural network on a large collection of gut microbiome samples to predict which species belong in a given community; it learns from co-occurrence during training but needs only gene annotations at inference. Working with a Sunagawa lab colleague, who built the model, and members of the von Mering lab, I designed and ran its validation, with co-occurrence, gene-content and phylogenetic baselines and an audit of the public training metadata that cleaned the training set.
  • Mentoring: mentored three PhD students and supervised a master's thesis (2025).

PhD researcher

Patil Lab, MRC Toxicology Unit, University of Cambridge · October 2020 to August 2024

  • Open-source software: built and released metaGEM (Nucleic Acids Research 2021, first author), a Snakemake workflow packaged on Bioconda, and ran it on HPC clusters to reconstruct 14,000+ genome-scale metabolic models from 483 metagenomes. Cited 150+ times and used by other groups in their own studies; six releases (2021 to 2023) and replies to 88 of the 90 GitHub issues users opened.
  • Soil cross-feeding (Nature Microbiology 2026, co-first author): co-designed the genome sequencing and modelling, and carried out most of the genomic analysis and metabolic modelling of 62 strains, compared with the Kost lab's growth measurements.
  • Collaborations: metabolic modelling for cheese flavour with Chr. Hansen (Nat. Commun. 2023); auxotrophy analysis for drug tolerance (Nat. Microbiol. 2022); metabolic models and metabolomics for C. difficile resistance (bioRxiv 2024); all metagenome assemblies and genome bins for a plastic-enzyme survey (mBio 2021).

Computational Biologist

Patil Lab, EMBL Heidelberg · August 2019 to July 2020

  • Started developing metaGEM, which became my PhD project.

Education

2020 to 2024
PhD, Biology · University of CambridgeMRC Toxicology Unit, Patil LabThesis: Omics-driven and constraint-based modelling of microbial community metabolism
2017 to 2019
MSc, Biotechnology · Chalmers University of TechnologyiGEM 2018, Chalmers-Gothenburg team: gold medal, and nominee for Best Model in the graduate section for COM-dFBA, the team's community dynamic flux balance analysis framework, which I ledElected treasurer of the Society for Biological Engineering students at Chalmers; managed events and budgets
2013 to 2017
BSc, Biological Systems Engineering · UC DavisSenior design project (team of three): a low-cost Arduino system for remote sensing of plant CO₂ uptake; I programmed the Arduino and assembled the sensor electronics

Skills

Bioinformatics: metaGEM, Snakemake, fastp, MEGAHIT, CONCOCT, MaxBin2, MetaBAT2, metaWRAP, GTDB-Tk, MMseqs2, DIAMOND

Systems biology: CarveMe, SMETANA, MEMOTE, COBRApy, ReFramed, RAVEN, FBA / FVA

Protein AI tools (pretrained models): ProstT5, Foldseek, AlphaFold2

Methods: model validation and benchmarking (baselines, cross-validation, bootstrap), pangenomics, phylogenetic regression, metabolomics

Languages and infrastructure: R, Python (NumPy, SciPy), MATLAB, Bash, Slurm HPC, GPU jobs on Slurm, Conda/Bioconda packaging, Git/GitHub, Claude Code (agentic coding)

Talks, teaching and supervision

2018 to 2026

I have taught in five courses (two EMBO Practical Courses, plus courses at EMBL-EBI, ETH Zürich and the University of Cambridge), with 75+ participants in total, and presented my research at conferences in Ireland, Germany and the US. Course materials are public.

Activity

Date Title Venue Role
2026 · Jan NCCR Microbiomes Winter Course: advanced methods in microbial community analysis UNIL · Lausanne prepared exercises
2025 · Nov Microbial Community Genomics (551-1119-00L) · Sunagawa Lab Block Course ETH Zürich instructor
2025 Master's thesis supervision (project hosted by a lab at the University of Oxford) ETH Zürich supervisor
2024 · Oct Metabolite and species dynamics in microbial communities EMBO Practical Course, Bangalore instructor
2024 · Oct Metabolic modelling for microbial ecology 9th COBRA Conference, San Diego poster
2024 · Jan Flux balance analysis and metabolic modelling Part III Systems Biology (master's level), Cambridge instructor
Earlier talks and courses (7, 2018 to 2022)
2022 · Oct Metabolic modelling of community interactions EMBO Practical Course (online) instructor
2022 · Oct Metagenomics-driven metabolic modeling for microbial ecology EMBO Workshop: Molecular mechanisms in evolution and ecology, Heidelberg poster and flash talk
2022 · Sep Metagenomics-driven metabolic modeling for microbial ecology 8th COBRA Conference, Galway selected talk
2022 · Jun Applications of genome scale metabolic models S2M2 Summer School in Metabolic Modelling, Braga (online) invited talk
2022 · Feb From metagenomics to metabolic interactions SymbNET 2022 Course, EMBL-EBI (online) instructor
2021 · Mar metaGEM: reconstruction of genome scale metabolic models directly from metagenomes 7th COBRA Conference (online) poster
2018 · Oct iGEM 2018 · Gold medal, Best Model nominee (graduate section) iGEM Giant Jamboree, Boston team member
Course materials (6)

Materials from these courses are public.

  • NCCR Microbiomes Winter Course
    2026 · UNIL Lausanne · Zenodo

    Materials for a course on advanced methods in microbial community analysis. I helped prepare the reproducible data analysis exercises.

  • Microbial Community Genomics (551-1119-00L)
    2025 · ETH Zürich · Fall block course

    Genome-scale metabolic models, community modelling and AI-based gene annotation. Tutor with Samuel Miravet-Verde and Martin Sperfeld; course run by Shinichi Sunagawa.

  • EMBOMicroCom2
    2024 · EMBO Practical Course, Bangalore

    Tutorial on flux balance analysis and genome-scale metabolic models.

  • systems-biology-fba-practical
    2024 · Part III Systems Biology, Cambridge

    Flux balance analysis practical, originally by Arianna Basile and Kiran Patil; I updated and taught it in 2024.

  • EMBOMicroCom
    2022 · EMBO Practical Course (online)

    Tutorial on flux balance analysis, genome-scale metabolic models and microbial ecology.

  • SymbNET
    2022 · EMBL-EBI (online)

    Walkthrough from metagenomes to community metabolic models.

Now

Autumn 2026
  • Manuscripts from ETH: two papers from my postdoc, on which I am co-first author, are heading to preprint. My co-leads on those modules are taking them forward, and I stay involved.

Side interests

  • Open-source sustainability: what keeps scientific software tools alive in the wild, and what kills them.
  • Low-cost lab hardware: cheap sensors and open code for phenotyping bacteria, yeast and fungi outside a big lab, from an Arduino CO₂ sensor at UC Davis to a home setup for growing gourmet mushrooms.
  • Music: amateur multi-instrumentalist.
  • Long-distance running: half-marathons and trail routes around Zürich.

Get in touch

Happy to talk about computational biology, open-source tools or a role on your team. Email is fastest.

Email
franciscozorrilla94@gmail.comBest way to reach me
LinkedIn
fzorrilla94Professional profile
GitHub
franciscozorrillaOpen-source work
Google Scholar
Francisco ZorrillaPublications and citations
Bluesky
@metagenomezOccasional posts and reposts
X
@metagenomezOccasional posts
Location
Zürich, Switzerland