Data on the web
- Summary
- Working with data on the Web is difficult due to numerous issues which an interested data consumer can come across, the main ones being data interoperability issues on various levels of abstraction. To support the ecosystem of data exchange on the Web, we are working on a set of techniques and tools for proper publishing and consumption of data on the Web, which include data cataloging, transformation, querying and visualization tools. We put data-centricity, data semantics and data distribution into the core of our work.
The Web is full of data represented in plenty of data formats and structures. It is hard for data consumers to find relevant data, integrate data from various data sources and interpret their meaning correctly. The naive old techniques of data centralization and unification of their syntax and semantics do not work any more. The data is too heterogeneous and constantly changing so these simple techniques are no longer applicable. This is not only true for the Web but also for the internal environment of every organization being small or large. Various studies show that up to 80 % of working with data takes finding, accessing and integrating the data and only 20% of valuable time of software and data engineers remains for creating added value on top of the data.
In our research, we develop novel techniques and tools which improve data quality. We focus on data findability, availability, interoperability and reusability (so called FAIR data principles). FAIR principles depend on properly described semantics of the data. It means that data consumers are able to easily find data they need using the meaning and to interpret the meaning of the data independently of technical formats and data structures used to represent the data.
We work on smart and easily usable techniques and tools for data transformation, publication, cataloging, integration and visualization. At their core, we put properly described and interpretable meaning of the data modeled as ontologies. We participate in the global movement to build a shared distributed Web of Data, historically called Semantic Web or Linked Data, today known under the term Knowledge Graphs. We also help private as well as public organizations to implement these techniques and tools to achieve FAIR principles in their own environment.
People
Latest publications
- Dataspecer: Development of Consistent Semantic Data Specification Ecosystems (2026)
- Improving Linked Data Development Experience with LDkit (2026)
- Data Specification Vocabulary (DSV): Representation of Application Profiles of Semantic Data Specifications (2025)
- Semantic web: Past, present, and future (2024)
- Enhancing domain modeling with pre-trained large language models: an automated assistant for domain modelers (2024)
- Towards Authoring of Vocabularies and Application Profiles using Dataspecer. (2024)
- Atlas: A toolset for efficient model-driven data exchange in data spaces (2023)
- LDkit: linked data object graph mapping toolkit for web applications (2023)
- Dataspecer: A model-driven approach to managing data specifications (2022)
- Open dataset discovery using context-enhanced similarity search (2022)
- Evaluation framework for search methods focused on dataset findability in open data catalogs (2020)
- Improving findability of open data beyond data catalogs (2019)
- Survey of tools for linked data consumption (2019)
- DCAT-AP representation of Czech National Open Data Catalog and its impact (2019)
- Advanced Analytics of Large Connected Data Based on Similarity Modeling (2018)
- UnifiedViews: An ETL tool for RDF data management (2018)
- Publication and usage of official Czech pension statistics Linked Open Data (2018)
- LinkedPipes ETL in use: practical publication and consumption of linked data (2017)
- LinkedPipes ETL: Evolved linked data preparation (2016)
- User Assisted Creation of Open-Linked Data for Training Web Information Extraction in a Social Network (2013)
Supporting software
Dataspecer
Dataspecer is a tool for the management and modeling of data structures. The structures may be mapped to user-provided ontologies and exported into various formats.
Knowledge Graph Browser
A tool for visual exploration of knowledge graphs through different views defined by various browsing configurations.
LinkedPipes ETL
LinkedPipes ETL is an open-source extract-transform-load tool focused on publication and consumption of data on the Web, including linked data. It is already deployed in multiple organizations world-wide and we have plenty of ideas of how to improve it further.
STIRdata
STIRdata is an EU project focused on improving interoperability of data about companies and organizations coming from business registries in individual european countries. In the project we apply our know-how in data modeling and data standardization.