deepFRI2 is an upgraded version of the well-established deepFRI (Deep Functional Residue Identification) framework for predicting protein function using Gene Ontology (GO) terms.
Like its predecessor, deepFRI2 operates in two complementary modes: sequence-based and sequence–structure-based. This dual approach enables robust functional inference in metagenomic settings — where protein structures are often unavailable — as well as structure-informed functional annotation when structural information is available.
For training, deepFRI2 leverages FRIData, a scalable and efficient library for generating large, non-redundant protein datasets.
While maintaining similar input and output formats, the model architecture has been completely redesigned to incorporate recent advances in the field, particularly the use of protein language models as powerful representations of protein sequences. It consists of two submodules: sequence analyzer (utilizing ESM embeddings and lightweight attention pooling) and structural prober (processing distograms with shallow convolutional network). Signals from both models are merged using an ESM-conditioned gating mechanism. deepFRI2 outputs sequence-, structure-, and fusion-based predictions for each ontology (MF, CC, BP). The architecture is intentionally simple yet robust, enabling accurate functional annotation while maintaining high scalability and interpretability.
condaormamba- For GPU inference: an NVIDIA GPU with a driver supporting CUDA 12.6
- For CPU inference: no additional requirements
# Clone the repository
git clone https://github.com/Tomasz-Lab/deepFRI2.git
cd deepFRI2
# Create the environment (choose ONE)
conda env create -f environment-gpu.yml # Recommended (GPU)
# conda env create -f environment-cpu.yml # CPU-only
# Activate the environment
conda activate deepfri2
# Download the ESM-2 language model (~2.5 GB)
python src/deepFRI2/download_esm.py- GPU environment is named
deepfri2, whereas CPU environment is nameddeepfri2_cpu(name can be changed in the corressponding.ymlfile). - The deepFRI2 model checkpoints are already included under
params/<ontology>/. Only the ESM-2 weights (downloaded in the last step) need to be fetched. - Once the ESM-2 weights are downloaded, all inference runs entirely offline.
Predict GO terms for a folder of protein structures (.cif / .pdb):
python src/deepFRI2/deepfri2.py --input path/to/structuresOr run the sequence model on a FASTA file of sequences (no structures needed):
python src/deepFRI2/deepfri2.py --input sequences.fastaOptions (run python src/deepFRI2/deepfri2.py --help for the full list):
| Flag | Description |
|---|---|
-i, --input |
Either a folder with .cif / .pdb structures or a single FASTA file of sequences. A FASTA input runs the sequence model only: --model and --ids_file are ignored, and --threshold uses only its first value. (required) |
-o, --output_dir |
Folder for results (default: <repo>/results). |
-f, --ids_file |
Text file listing structures to run, one per line. Each entry is an id, optionally with a .cif / .pdb extension and/or a relative subfolder path (abCD, abCD.cif, sub/abCD, sub1/sub2/abCD.cif), resolved under the input folder. An id without an extension resolves to .cif if present, else .pdb. Default: all top-level files in the input folder. |
-a, --aspect |
Comma-separated GO aspects (ontologies) to run: any of MF, CC, BP (case-insensitive). Default: mf,cc,bp. |
-m, --model |
Which model to run: sequence (embeddings only), structure (distograms only) or fusion (both). Only the needed inputs are computed and outputs carry only that model's columns. Default: fusion. |
-b, --batch_size |
Proteins per inference batch (default: 32). |
-t, --threshold |
Keep a GO term in the summary if any model (sequence, structure, fusion) scores ≥ threshold. Either one float applied to all models (0.1) or two comma-separated floats applied to fusion/sequence and structure respectively (0.1,0.2). The structural prober is trained with a different loss and outputs higher probabilities on average, hence the higher default for it. 0 (or 0,0) keeps everything; 1 (or 1,1) keeps nothing (default: 0.1,0.2). |
-k, --top_k |
Maximum GO terms per protein in the summary (default: all selected). |
-p, --prop |
Propagate scores up the GO hierarchy. Adds the preds_propagated/ folder and the propagated columns to the summary (default: off). |
-s, --summary |
Write only prediction_summary.csv, skipping the preds/ (and preds_propagated/) folders (default: off). With --prop, the summary still includes the propagated columns. |
-v, --verbose |
Enable debug logging (default: off). |
Run a subset of structures with a stricter (global) thresholds:
python src/deepFRI2/deepfri2.py -i structures/ -o results/run1 -f ids.txt -t 0.3- GPU environment is the default and recommended one.
- The model runs in float32. GO ranking and thresholded term sets are effectively identical across CPU/GPU. Differences in probabilities are of order
$\mathcal{O}(10^{-7})$ between GPUs (the same CUDA versions) and$\mathcal{O}(10^{-6})$ between GPU and CPU. - In the current setup, the model processes up to 1,020 aa. Longer proteins are truncated during inference (no functional signal beyond 1,020 aa), but ESM embeddings are generated for the full protein, which may be time-consuming. Therefore, it is recommended to run the model on functionally relevant domains. There is no lower limit; however, the structural prober is not sensitive to proteins shorter than 60 aa (in which case predictions equal the mean across the training data).
- The current version was trained on gapless structures, so fully resolved inputs (no missing residues) are recommended. For structures with gaps, a missing residue zeroes out its whole row/column in the residue–residue similarity map, breaking the backbone-adjacency band and pushing the structure model out of distribution. As a rough safeguard we fill only the immediate
-1/+1off-diagonals at gap positions with1(a non-zero, "these consecutive residues are neighbours" signal).
Predictions are written under the output folder for ontologies selected by the --aspect argument (MF, CC, BP by default):
prediction_summary.csv— top predicted GO terms per protein, with raw scores for the fusion, structure, and sequence models (and GO-hierarchy-propagated scores when--propis set).preds/<protein>__<ontology>.csv— full per-term probabilities (fusion / structure / sequence + gate). Omitted when--summaryis set.preds_propagated/<protein>__<ontology>.csv— full per-term probabilities after GO-hierarchy propagation (only when--propis set; omitted when--summaryis set).log.txt— the run log.
For a quick overview of predicted functions, please take a look at the pred_prob (raw probabilities) column in prediction_summary.csv — or, when you run with --prop, the pred_prop_prob column (consistent probabilities i.e., the more general the term, the higher its probability). In some cases, it is also useful to check purely structure- and sequence-based outputs (see struct_prob, seq_prob etc.). For a downstream analysis, you may wish to check the full output in preds (and, with --prop, preds_propagated) folders.
End-to-end runtime, excluding model loading at startup (which typically takes a couple of seconds per run), depends primarily on the available computational resources and protein length. Initial benchmarks using the default batch size of 32 yielded the following throughput:
| GPU (NVIDIA A100) | CPU server (48 cores) | Laptop (Apple M3 Pro, 11 cores) | |
|---|---|---|---|
| Fusion | 0.2–0.3 | 0.6–1.3 | 33 |
| Sequence | 0.1–0.2 | 0.4–1.0 | 2.3–6.2 |
These measurements, with the exception of the laptop benchmark for the fusion model (where a dataset with shorter proteins was used), were obtained on protein datasets with median sequence lengths ranging from 150 to 440 amino acids.
For large-scale inference, we recommend using a GPU or a multi-core CPU server. On CPU, ESM embeddings are computed one sequence at a time, with each forward pass utilizing all available cores.
The model can also be run on a personal computer, such as a laptop. The structure model is considerably slower on Apple Silicon because the PyTorch CPU build does not use MKL. Local CPU inference is therefore best suited for small runs (e.g., by selecting a subset of .cif/.pdb files with --ids_file) or for sequence-only inference. In the latter case, most of the computation time is spent generating embeddings (approximately 97–99%), while the actual model inference is very fast.
The model is still under development. We will soon add (among other things):
- interpretability module
- architectural details
- detailed benchmarks
In the nearest future we also plan to share the whole training pipeline in a fully reproducible manner.
If you run into installation problems, find a bug, or would like to propose an improvement, please raise an issue or write directly to p.szczerbiak[at]sanoscience.org.
