F. Zorrilla

About

Looking for my next role

Roles
Senior, Lead or Principal positions in computational biology, bioinformatics or scientific data
Sectors
Biotech, pharma, food or agtech
Location
Zürich area preferred, open to roles anywhere in Switzerland
Start
From December 2026, flexible on the exact date
Permit
Italian citizen with a Swiss B permit, no sponsorship needed
Languages
English and Spanish (native level), Italian (B1), French (A1), German (A1)
Skills
Metagenomics and microbial genome analysis · genome-scale metabolic modelling · reproducible pipelines (Snakemake, Conda, Slurm) · Python and R · protein language models and structure search (ProstT5, Foldseek)

What I bring: an open-source workflow I built and support (metaGEM); metabolic models built with wet-lab partners and checked against their measurements (co-first author, Nature Microbiology 2026, with the Kost lab) or used alongside their experiments (Nature Communications 2023, with Chr. Hansen); and teaching five courses.

How I could help an R&D team

  • Predict from genomes which nutrients the strains in a collection cannot make, and check those predictions against growth data, including where they fail (Nature Microbiology 2026).
  • Model how the strains in a starter culture or consortium feed each other, alongside fermentation and omics data (Nature Communications 2023, with Chr. Hansen).
  • Build a draft genome-scale metabolic model of a production organism from its genome (Hansenula protocol, 2022).
  • Annotate genes that sequence similarity misses, using protein language models and structure search (ProstT5, Foldseek), and feed the results into metabolic models.
  • Build reproducible analysis pipelines that other groups install and run on their own data (metaGEM).
  • Turn a broad research goal into a plan with clear modules, bring the right people in to lead them, and keep them on track (NCCR Microbiomes flagship, six labs).

Overview

I have worked in metagenomics and metabolic modelling for 7+ years, and since 2024 also with protein language models and structure search for gene annotation. I joined the Patil Lab at EMBL Heidelberg in 2019 and moved with the group to the MRC Toxicology Unit, University of Cambridge, where I did my PhD. From 2024 to 2026 I was a postdoc in the Sunagawa Lab at ETH Zürich, where I was co-lead of the NCCR Microbiomes Work Package 5 flagship project (details in the CV). My work turns raw metagenomic data into predictions of what each microbe in a community eats and makes, and I check those predictions against experimental data from partner labs: the Kost lab (Osnabrück) on soil isolates, Joy Scaria's lab on C. difficile bioreactors, and Chr. Hansen on cheese starter cultures.

Code, models and data from metaGEM, the soil and cheese papers and my PhD thesis are public on GitHub and Zenodo.

Work

12 publications

8 journal papers, 1 book chapter, 2 preprints and my PhD thesis. Tap a card to see the problem, the approach, my contribution and links. Citations from Google Scholar, news and social counts from Altmetric, as of Sept 2026.

Browse by collaborator

Collaboration map: 12 publications and 13 key co-authors. Tap or hover a dot for details; tap a paper to open its summary.

Co-author Paper Industry collab

Highlights

Sort:

Problem

Going from metagenomes to community metabolic models meant chaining a dozen separate tools by hand, and models built from reference genomes miss the strain-level differences in a real sample.

Approach

I designed and released metaGEM, a Snakemake workflow that runs quality control, assembly, binning, bin refinement, taxonomy, metabolic model reconstruction and community simulation in one workflow, on a single machine or a Slurm cluster.

My contribution

First author and lead developer. The paper credits me with co-conceiving the project, developing the pipeline with Filip Buric, running all the analyses and co-writing the paper. I wrote the documentation, maintain the Bioconda package, published six releases (v1.0.0 to v1.0.5, 2021 to 2023) and still answer users' GitHub issues.

Tools

Snakemake, Python, CarveMe, SMETANA, MEGAHIT, CONCOCT, MaxBin2, MetaBAT2, metaWRAP, GTDB-Tk, MEMOTE, Slurm HPC

Links

Problem

Ecological theory predicts that many soil bacteria cannot make some essential nutrients and rely on neighbours for them, which would also help explain why most soil bacteria will not grow alone in the lab. This had not been tested systematically in real soil communities. For agtech, a strain that relies on neighbours for nutrients may not establish on its own.

Approach

The Kost lab isolated 6,931 bacteria from 27 soil communities in Germany and tested which could grow without added nutrients. We sequenced 62 strains and built metabolic models for them to see whether their genomes predict these dependencies. The lab then grew auxotrophs with strains from the same soil, in pairs and small groups, and the models predicted which neighbours could supply the missing nutrients.

My contribution

Co-first author, shared with Ghada Yousif. I co-designed the genome sequencing and modelling and did most of the genomic analysis and metabolic modelling: sequence analysis with metaGEM, genome-scale models of the 62 sequenced strains and their comparison with the Kost lab's growth tests, the insertion-sequence and gene-loss analysis, pangenomics, the phylogenetic regression and the community simulations. I also co-wrote the first draft. The auxotrophy screen was the Kost lab's experimental work. The models missed most auxotrophies, and the genome analysis linked many missed cases to insertion sequences and gene loss in pathway genes, which became a result of the paper. The code repository maps every figure and table to the code and data that produce it.

What I used

metaGEM, CarveMe, SMETANA, ReFramed, Pangenomics, R (phyloglm), Python

Press coverage

Links

Problem

Cheddar starter cultures mix Streptococcus thermophilus with several Lactococcus strains. How these strains interact during ripening, and how that shapes flavour, was largely unknown beyond pairwise tests in simplified lab media.

Approach

The team, led by first author Chrats Melkonian (Chr. Hansen and Wageningen University), made Cheddar with the full commercial starter culture and with single strains left out, and followed it through a year of ripening; shorter milk fermentations left out single Lactococcus strains. They combined genome sequencing, genome-scale metabolic models, metatranscriptomics and metabolite measurements. S. thermophilus breaks down milk protein, which relieves nitrogen limitation for Lactococcus, and a Lactococcus cremoris strain keeps the off-flavour compounds diacetyl and acetoin in check.

My contribution

Second author. I did most of the published metabolic modelling, together with first author Chrats Melkonian: genome-scale metabolic models of the starter strains built with metaGEM's modelling module (CarveMe, SMETANA), and the community simulations. This built on earlier modelling by Daniel Machado. I also helped write the paper. The cheese-making experiments were run at Chr. Hansen.

What I used

metaGEM, CarveMe, SMETANA, ReFramed, COBRApy, Python

Press coverage

Links

Problem

Recurrent C. difficile infection responds to faecal transplant, but the active ingredients of "colonisation resistance" are unclear. Which metabolic interactions actually keep the pathogen out?

Approach

Joy Scaria's lab (Oklahoma State and South Dakota State universities) challenged a 14-strain gut community, and communities grown from stool samples, with antibiotics and then C. difficile in a continuous-flow bioreactor; the Patil lab in Cambridge did the modelling. Population models and genome-scale metabolic models pointed to interactions between E. coli and Bacteroides/Phocaeicola, and metabolomics pointed to fructooligosaccharide use, vitamin B3 synthesis and competition for the amino acids C. difficile ferments. Patient metagenomes supported these signals.

My contribution

Third author. With Arianna Basile I annotated the genomes and built and simulated genome-scale metabolic models of the community members and C. difficile (CarveMe, SMETANA). I analysed the metabolomics data comparing communities that suppressed C. difficile with those that did not, checked the genomes for the genes behind those differences, and helped analyse the patient metagenomes.

What I used

CarveMe, KEMET, SMETANA, ReFramed, R

Links

Problem

Building and curating a genome-scale metabolic model is slow and tedious, and Hansenula polymorpha (Ogataea polymorpha), an industrially relevant methylotrophic yeast, had no published high-quality model.

Approach

A step-by-step protocol chapter that builds the model from the models of related yeasts with the RAVEN toolbox (MATLAB), and releases the result as hanpo-GEM.

My contribution

First author, with Eduard Kerkhoven (Chalmers). We built the draft model together as a side project during my MSc, and I wrote the protocol chapter; I published the model's v1.0 release on GitHub in 2024.

Tools

RAVEN, MATLAB, SBML, Homology-based reconstruction

Links

Problem

Plastic pollution is growing faster than the list of enzymes known to break plastic down, and finding new ones in the lab is slow. The question was whether global metagenomic data could point to candidates.

Approach

The team (first author Jan Zrimec, Chalmers) built profile hidden Markov models from 95 experimentally verified plastic-degrading enzymes and searched ocean and soil metagenomes from 236 sites, finding over 30,000 candidate enzymes for 10 plastic types. Enzyme counts correlated with measured ocean plastic pollution and, for soils, with country-level data on poorly managed plastic waste.

My contribution

Fourth author. I generated all the metagenome assemblies and genome bins (MAGs) for the study with metaGEM; Jan Zrimec then ran the enzyme searches and all the statistics.

What I used

metaGEM, MEGAHIT, CONCOCT, MetaBAT2, MaxBin2, metaWRAP

Press coverage

Links

Other research (5 papers)
  • Drug tolerance via cross-feeding

    Nat. Microbiol. 2022 · Third author of 19 · 144 citations · Paper (Nat. Microbiol.)

    I did the analysis that opens the paper: how common amino-acid auxotrophs are across more than 12,000 Earth Microbiome Project communities, and whether auxotrophs grow better under drugs.

  • Enterobacteriaceae gut dynamics

    Nat. Microbiol. 2025 · Third author · 97 citations · Paper (Nat. Microbiol.)

    I gave the authors a session on flux balance analysis and genome-scale metabolic models, then reviewed the finished paper and helped interpret the metabolic scores.

  • FROG: GEM reproducibility

    bioRxiv 2024 · Co-author · 15 citations · bioRxiv preprint

    As a community contribution, I submitted a genome-scale model from one of our lab's published studies to BioModels as a test case.

  • iGEM interlab measurement studies

    Commun. Biol. 2020 and PLOS ONE 2021 · Consortium author · 381 and 16 citations · Paper (Commun. Biol.) · Paper (PLOS ONE)

    As a master's student on the Chalmers iGEM 2018 team, I was one of three members in charge of our lab's interlab contribution: growing the test strains, calibrating OD and fluorescence on the plate reader, and counting colonies.

  • PhD thesis

    University of Cambridge 2024 · Full thesis (open access) · GitHub

    Omics-driven and constraint-based modelling of microbial community metabolism.

metaGEM

Open-source software

Builds genome-scale metabolic models directly from metagenomes.

What it does

metaGEM is an end-to-end Snakemake workflow that turns raw metagenomic reads into community-level metabolic predictions.

Pipeline stages:

  1. QC (fastp)
  2. Assembly (MEGAHIT)
  3. Binning (CONCOCT, MaxBin2, MetaBAT2)
  4. Bin refinement and reassembly (metaWRAP)
  5. Taxonomy (GTDB-Tk)
  6. Metabolic reconstruction (CarveMe)
  7. Model QC (MEMOTE)
  8. Community simulation (SMETANA)

Why this matters for an industry team: people outside my group install metaGEM from Bioconda and run it on their own data, and I have supported them since 2021, through six releases and replies to 88 of the 90 GitHub issues other users opened.

Install

git clone https://github.com/franciscozorrilla/metaGEM.git
cd metaGEM/workflow
mamba env create -n metagem -f envs/metaGEM_env.yml

Linux (clusters or workstations). metaWRAP, the modelling tools (CarveMe, MEMOTE, SMETANA) and the reference databases need a few more steps: see the setup guide.

Worked example: metaGEM on Unseen Bio gut samples

Applications in the wild

Published studies in which other groups used metaGEM, or followed its modelling steps, on their own data. Each one has been checked against the paper's methods section.

  • 2026 · Nature Communications

    Continental-scale soil carbon decomposition

    Song et al. (Pacific Northwest National Laboratory, the DOE Joint Genome Institute and Eawag) used metaGEM for quality control, assembly and binning of metagenomes from 47 US soil cores, then linked the recovered genomes to soil organic-matter chemistry.

    Carbon cycling · climate

  • 2026 · Cell Host & Microbe

    Soil protists and bacterial cooperation

    Liu et al. (Nanjing Agricultural University) extended metaGEM with an extra binning tool (SemiBin) to build community metabolic models from rhizosphere metagenomes. These models confirmed their main finding: predation by soil protists shifts bacteria from competing to cooperating.

    Soil ecology · agtech

  • 2025 · Environmental Microbiology Reports

    Subsurface H₂ storage microbiology

    Tinker et al. (US national laboratories NETL, PNNL and Sandia) ran metaGEM with default settings to recover genomes from water in a deep saline aquifer in Illinois, to assess how storing hydrogen underground could affect its microbes.

    Energy · underground hydrogen storage

  • 2024 · Communications Biology

    Microbiome for industrial phenolic wastewater

    Zhao et al. (Tianjin Institute of Industrial Biotechnology, CAS) used metaGEM to recover 164 genomes from a microbial community they scaled up from shake flasks to an industrial plant treating phenolic resin wastewater.

    Industrial wastewater · biotech

Show 2 more
  • 2022 · F1000Research

    NEON soil metagenomes pipeline

    Werbin et al. (Boston University) moved the main workflow of their public tutorial for the National Ecological Observatory Network's soil metagenomes to metaGEM, citing its support for computing clusters, and thanked the metaGEM developers for help troubleshooting.

    Soil · large-scale ecology

  • 2022 · Environmental Microbiome

    Movile Cave: chemoautotrophic ecosystem

    Chiciudean et al. (Babeș-Bolyai University, Romania) followed metaGEM's modelling steps (CarveMe, MEMOTE, SMETANA) on their own genomes from a sulfidic cave, to map competition and cooperation in an ecosystem that runs without sunlight. I helped with their community simulations.

    Extreme environments · ecology

Selected from the 150+ works that cite metaGEM. See the full list on Google Scholar.

Curriculum vitae

Postdoc, ETH Zürich (2024 to 2026) · PhD, University of Cambridge · EMBL Heidelberg

Download CV (PDF)

Experience

Postdoctoral Researcher · Co-lead, NCCR Microbiomes Work Package 5 Flagship Project

ETH Zürich · October 2024 to September 2026

  • Planning: designed the 4-year computational roadmap and its modules for the flagship (2024 to 2028; ETH Zürich, EPFL, UZH, UNIL and CHUV) and pitched it to the PIs.
  • Leadership: coordinated the flagship with a PhD co-lead from EPFL on behalf of the work-package leaders. I organised its launch workshop (Bern, March 2025), recruited the co-leads of the four modules, ran about 20 coordination meetings and the progress reports, and co-led two modules until September 2026, with 50+ working sessions. I am co-first author on both module manuscripts (in preparation).
  • Module 4, structure-based metabolic models (with EPFL): chose ProstT5 and Foldseek over AlphaFold so annotation scales to microbiome gene sets; built the GPU annotation pipeline (Snakemake, Slurm) with custom enzyme and transporter databases; designed the comparison against sequence search (MMseqs2, DIAMOND) and experimental growth phenotypes.
  • Module 3, machine learning for community assembly: designed and ran the validation of a neural network that predicts gut community membership from gene content, with co-occurrence, gene-content and phylogenetic baselines, fold-aware statistics, and an audit of the public training metadata that cleaned the training set.
  • Mentoring: mentored three PhD students and supervised a master's thesis (2025).
  • Tools: since spring 2026 I have used Claude Code daily to prototype tools and extend analyses, and I check its output before relying on it.

PhD researcher

Patil Lab, MRC Toxicology Unit, University of Cambridge · October 2020 to August 2024

  • Open-source software: built and released metaGEM (Nucleic Acids Research 2021, first author), a Snakemake workflow packaged on Bioconda, and ran it on HPC clusters to reconstruct 14,000+ genome-scale metabolic models from 483 metagenomes. Six releases (2021 to 2023); I still answer users' GitHub issues.
  • Soil cross-feeding (Nature Microbiology 2026, co-first author): co-designed the genome sequencing and modelling, and did most of the genomic analysis and metabolic modelling of 62 strains, compared with the Kost lab's growth measurements.
  • Collaborations: metabolic modelling for cheese flavour with Chr. Hansen (Nat. Commun. 2023); auxotrophy analysis for drug tolerance (Nat. Microbiol. 2022); metabolic models and metabolomics for C. difficile resistance (bioRxiv 2024); all metagenome assemblies and genome bins for a plastic-enzyme survey (mBio 2021); advice on metabolic modelling for gut Enterobacteriaceae (Nat. Microbiol. 2025).

Computational Biologist

Patil Lab, EMBL Heidelberg · August 2019 to July 2020

  • Started developing metaGEM, which became my PhD project.

Education

2020 to 2024
PhD, Biology · University of CambridgeMRC Toxicology Unit, Patil LabThesis: Omics-driven and constraint-based modelling of microbial community metabolism
2017 to 2019
MSc, Biotechnology · Chalmers University of TechnologyiGEM 2018, Chalmers-Gothenburg team: gold medal, and nominee for Best Model in the graduate section for COM-dFBA, the team's community dynamic flux balance analysis framework, which I ledElected treasurer of the Society for Biological Engineering students at Chalmers; managed events and budgets
2013 to 2017
BSc, Biological Systems Engineering · UC DavisSenior design project (team of three): a low-cost Arduino system for remote sensing of plant CO₂ uptake; I did the Arduino coding and wiring

Skills

Bioinformatics: metaGEM, Snakemake, fastp, MEGAHIT, CONCOCT, MaxBin2, MetaBAT2, metaWRAP, GTDB-Tk

Systems biology: CarveMe, SMETANA, MEMOTE, COBRA, FBA / FVA, RAVEN

Protein AI tools (pretrained models): ProstT5, Foldseek, AlphaFold2

Languages and infrastructure: Python, R, MATLAB, Bash, Slurm HPC, GPU jobs on Slurm, Conda/Bioconda packaging, Git/GitHub, Claude Code (agentic coding)

Talks, teaching and supervision

2018 to 2026

I have taught in five courses (two EMBO Practical Courses, plus courses at EMBL-EBI, ETH Zürich and the University of Cambridge), with 75+ participants in total, and presented my research at conferences in Ireland, Germany and the US. Course materials are public.

Activity

Date Title Venue Role
2026 · Jan NCCR Microbiomes Winter Course: advanced methods in microbial community analysis UNIL · Lausanne prepared exercises
2025 · Nov Microbial Community Genomics (551-1119-00L) · Sunagawa Lab Block Course ETH Zürich instructor
2025 Master's thesis supervision (project hosted by a lab at the University of Oxford) ETH Zürich supervisor
2024 · Oct Metabolite and species dynamics in microbial communities EMBO Practical Course, Bangalore instructor
2024 · Oct Metabolic modelling for microbial ecology 9th COBRA Conference, San Diego poster
2024 · Jan Flux balance analysis and metabolic modelling Part III Systems Biology (master's level), Cambridge instructor
2022 · Oct Metabolic modelling of community interactions EMBO Practical Course (online) instructor
2022 · Oct Metagenomics-driven metabolic modeling for microbial ecology EMBO Workshop: Molecular mechanisms in evolution and ecology, Heidelberg poster and flash talk
2022 · Sep Metagenomics-driven metabolic modeling for microbial ecology 8th COBRA Conference, Galway selected talk
2022 · Jun Applications of genome scale metabolic models S2M2 Summer School in Metabolic Modelling, Braga (online) invited talk
2022 · Feb From metagenomics to metabolic interactions SymbNET 2022 Course, EMBL-EBI (online) instructor
2021 · Mar metaGEM: reconstruction of genome scale metabolic models directly from metagenomes 7th COBRA Conference (online) poster
2018 · Oct iGEM 2018 · Gold medal, Best Model nominee (graduate section) iGEM Giant Jamboree, Boston team member
Course materials (6)

Materials from these courses are public.

  • NCCR Microbiomes Winter Course
    2026 · UNIL Lausanne · Zenodo

    Materials for a course on advanced methods in microbial community analysis. I helped prepare the reproducible data analysis exercises.

  • Microbial Community Genomics (551-1119-00L)
    2025 · ETH Zürich · Fall block course

    Genome-scale metabolic models, community modelling and AI-based gene annotation. Tutor with Samuel Miravet-Verde and Martin Sperfeld; course run by Shinichi Sunagawa.

  • EMBOMicroCom2
    2024 · EMBO Practical Course, Bangalore

    Tutorial on flux balance analysis and genome-scale metabolic models.

  • systems-biology-fba-practical
    2024 · Part III Systems Biology, Cambridge

    Flux balance analysis practical, originally by Arianna Basile and Kiran Patil; I updated and taught it in 2024.

  • EMBOMicroCom
    2022 · EMBO Practical Course (online)

    Tutorial on flux balance analysis, genome-scale metabolic models and microbial ecology.

  • SymbNET
    2022 · EMBL-EBI (online)

    Walkthrough from metagenomes to community metabolic models.

Now

Autumn 2026
  • Manuscripts from ETH: two papers from my postdoc, on which I am co-first author, are heading to preprint. My co-leads on those modules are taking them forward, and I stay involved.
  • Looking for my next role from December 2026. Details →

Side interests

  • Open-source sustainability: what keeps scientific software tools alive in the wild, and what kills them.
  • Low-cost lab hardware: cheap sensors and open code for phenotyping bacteria, yeast and fungi outside a big lab. It started with a low-cost Arduino sensor for plant CO₂ uptake (my team's senior design project at UC Davis; I did the coding and wiring) and continued with a home setup for growing gourmet mushrooms.
  • Music: amateur multi-instrumentalist.
  • Long-distance running: half-marathons and trail routes around Zürich.
  • International background: I have lived in 11 countries across the Americas, Europe and Asia.

Get in touch

Happy to talk about computational biology, open-source tools or a role on your team. Email is fastest.

Email
franciscozorrilla94@gmail.comBest way to reach me
LinkedIn
fzorrilla94Professional profile
GitHub
franciscozorrillaOpen-source work
Google Scholar
Francisco ZorrillaPublications and citations
Bluesky
@metagenomezOccasional posts and reposts
X
@metagenomezOccasional posts
Location
Zürich, Switzerland

Key numbers