Multimedia Retrieval
- Summary
- Multimedia data penetrate all areas of our lives and become more important than ever. We meet them in social media applications, video streaming services, digital libraries as well as in specialized medical or industrial fields. As multimedia data are produced using sensors, their primary representation is semantically unstructured (e.g., an image is a bunch of pixels). Hence, recognition of what actually is inside a particular multimedia document and subsequent retrieval is a hard task that requires advanced techniques for feature extraction, object detection, similarity modeling, etc. Many of these techniques are based on machine-learning models. We carry out research in various multimedia retrieval problems and also propose many topics for student academic works (Bc, Mgr, PhD).
When working with data and databases, we traditionally consider some kind of structured data. Tabular data (relational databases, spreadsheets), XML/JSON files, graph data (RDF) — these all are data formats with entities structured into attributes. For them a schema is usually provided followed by a specification/query language on how to access such data. On the other hand, multimedia data is a special kind of data that is not structured into meaningful attributes.
To discover a structure and semantics in multimedia data, we need to employ specific methods for feature extraction and/or object detection. Extracted features then form more or less semantic descriptors that could be used for content-based multimedia retrieval. The content descriptors are black boxes to the users, so their internals (attributes) cannot be used by them directly to construct queries. Instead, a similarity function is provided along the descriptors that is designed to aggregate the differences between two descriptors (original objects, respectively) into a real-value similarity score. For queries the similarity function is applied in the query-by-example fashion on a query object (the example) and the database objects. A ranking based on the similarity scores is established and the best results are returned to the user.
People
Tomáš Skopal
Deputy head of department, Professor
Jakub Lokoč
Associate professor
Ladislav Peška
Associate professor
Latest publications
- What Drove Success at the 15th Video Browser Showdown? (2026)
- Evaluating Keyframe Layouts for Visual Known-Item Search in Homogeneous Collections (2026)
- CALISSTA: Calibrated Candidate Retrieval Using Sparse Autoencoders and Top-k Aggregation (2026)
- VISAnt: Unsupervised Data Exploration with Chernoff Faces (2025)
- Results of the 2025 Video Browser Showdown (2025)
- Unified visual-aware representations for data analytics (2025)
- RESET: Relational Similarity Extension for V3C1 Video Dataset (2024)
- Evaluating performance and trends in interactive video retrieval: Insights from the 12th vbs competition (2024)
- Visualizations for universal deep-feature representations: survey and taxonomy (2024)
- Interactive multimodal video search: an extended post-evaluation for the VBS 2022 competition (2024)
- Prak tool: An interactive search tool based on video data services (2024)
- Known-item search in video: An eye tracking-based study (2024)
- Introduction to the seventh annual lifelog search challenge, lsc'24 (2024)
- Less is more: similarity models for content-based video retrieval (2023)
- Marine video kit: a new marine video dataset for content-based analysis and retrieval (2023)
- Comparing interactive retrieval approaches at the lifelog search challenge 2021 (2023)
- SpotifyExplained: User-centric Mobile Application for Music Exploration (2023)
- Visual representations for data analytics: user study (2023)
- Evaluating a Bayesian-like relevance feedback model with text-to-image search initialization (2023)
- Video search with CLIP and interactive text query reformulation (2023)