Bioinformatics
- Summary
- Bioinformatics sits on the interface of computer science and molecular biology where it helps to manage, understand and analyze biological and biomedical data. Our research in bioinformatics focuses on the development of software tools applicable mainly in the domain of structural bioinformatics and visualization. These include tools for protein binding site detection, with the application in computational drug discovery, or tools for visualization of the structure of macromolecules. All our methods are implemented as software solutions used by thousands of users all over the world. Most of our tools were first implemented as bachelor or master thesis in our department.
All living forms are based on the interaction of different types of molecules. At the center of bioinformatics are three types of macromolecules: DNA, RNA, and proteins. These, together with other molecular players, such as ions, small molecules, and lipids, interact, forming complex living systems.
Each of the macromolecules can be computationally represented depending on the level of abstraction as primary, secondary, or tertiary structure. As each of the three types of molecules consists of a linear chain of building blocks, so-called residues (nucleotides in case of DNA and RNA, and amino acids in case of proteins) can be represented as strings over a given alphabet. This representation is called primary structure. However, the residues fold in three-dimensional space into complex shapes. If we know which residues are close to each other (not necessarily knowing their exact location) we talk about secondary structure. If the 3D positions of all atoms are available, we talk about tertiary/quaternary structure.
Bioinformatics is divided into two branches: sequential and structural bioinformatics. Sequential bioinformatics consists of algorithms that work mainly over the primary structure of a molecule. On the other hand, algorithms that operate over the secondary/tertiary/quaternary structures are part of structural bioinformatics. Obviously, the areas are not strictly divided.
In our group, we mainly focus on the development of structural bioinformatics software solutions. Currently, our focus is to apply various machine learning techniques to detect different types of active sites on the surface of the protein structure. In another branch of research, we focus on the visualization of different types of molecules such as RNA secondary structures or protein tertiary structures. Here we both develop algorithms rooted in graph theory and software tools such as web servers or web-based visualization plugins.
We are members of various international consortia such PDBe-KB or RNAcentral which ensures that our software solutions are used by thousands of users all over the world.
Our interest in the bioinformatics also gave rise to the bioinformatics study program, which we co-created with our colleagues from the Faculty of Science. This cooperation also led to the formation of the Charles University Structure Bioinformatics Group.
People
David Hoksza
Head of department, Associate professor
Petr Škoda
Researcher
Samuel Fanči
PhD student
Hamza Gamouh
PhD student
Yana Podlesna
PhD student
Kamila Reimitz Riedlová
Researcher
Vít Škrhák
PhD student
Latest publications
- Multi-omic data fusion reveals the in vivo enzyme kinetics of Vibrio natriegens at the genome-scale (2026)
- Protein Language Models and Structure-Based Machine Learning for Prediction of Allosteric Binding Sites in Protein Kinases: An Explainable AI Framework Grounded in Energy Landscape-Encoded Frustration (2026)
- Predicting and Decoding Allosteric Binding Sites Using Protein Language Models and Structure-Based Machine Learning: An Energy Landscape-Guided Explainable AI Framework (2026)
- Beyond Exact Matches: Near-Hit Scoring for Protein Binding Site Prediction with Protein Language Models (2025)
- Hidden in protein sequences: Predicting cryptic binding sites (2025)
- CryptoBench: cryptic protein--ligand binding sites dataset and benchmark (2025)
- PrankWeb 4: a modular web server for protein--ligand binding site prediction and downstream analysis (2025)
- COBREXA 2: tidy and scalable construction of complex metabolic models (2025)
- Algebraic differentiation for fast sensitivity analysis of optimal flux modes in metabolic models (2025)
- Genomics 2 Proteins portal: a resource and discovery tool for linking genetic screening outputs to protein sequences and structures (2024)
- Cryptic binding site prediction with protein language models (2023)
- Interrogating the effect of enzyme kinetics on metabolism using differentiable constraint-based models (2022)
- COBREXA.jl: constraint-based reconstruction and exascale analysis (2022)
- PrankWeb 3: accelerated ligand-binding site predictions for experimental and modelled protein structures (2022)
- R2DT is a framework for predicting and visualising RNA secondary structure using templates (2021)
- PDBe-KB: a community-driven resource for structural and functional annotations (2020)
- Comprehensive characterization of amino acid positions in protein structures reveals molecular effect of missense variants (2020)
- Generalized EmbedSOM on quadtree-structured self-organizing maps (2019)
- PrankWeb: a web server for ligand binding site prediction and visualization (2019)
- MINERVA API and plugins: opening molecular network analysis and visualization to the community (2019)
Supporting software
COBREXA.jl
COBREXA.jl is an HPC-capable metabolic modeling toolbox, with advanced functionality for simulating microbiomes and bacterial communities.
MolArt
MolArt is a responsive, easy-to-use JavaScript plugin which enables users to view annotated protein sequence and overlay the annotations over a corresponding experimental or predicted protein structure.
P2Rank
P2Rank is a state-of-the-art machine learning-based method for ligand binding sites prediction based on protein structure.
Traveler
Traveler is an RNA sescondary structure visualization tool implementing a template-based approach enabling to lay out even the largest RNA structures in the standard orientation.