Focus

Data Science

Articles, podcasts, talks, and more about Data Science.
Blog Post

The Data Dilemma, Part 2

The first part showed why agents fail on undocumented data semantics and why Data Mesh is not a realistic answer for many mid-sized companies. This second part gets concrete: a Minimum Viable Data Mesh in four layers, from an honest inventory through curated views with ODCS contracts to agent access via MCP. A purchasing agent that compares suppliers and evaluates price histories serves as a running example.

Blog Post

The Data Dilemma, Part 1

When AI agents fail in a company, it is rarely the model’s fault. The real problem is the data: semantics, quality, and context are written down nowhere, they live in the heads of the domain experts. Data Mesh addresses exactly this problem, but it organizationally overwhelms many mid-sized companies. This first part describes why that is and derives a Minimum Viable Data Mesh from it. The second part shows the implementation in four layers.

Blog Post

The Missing Half of Your Data Strategy

Why Data Literacy Makes or Breaks Data Product Oriented Architectures

Blog Post

Data Inventories in the EU Data Act: The Democratization of IoT Devices

Starting in September 2025, the EU Data Act (Regulation (EU) 2023/2854) will require companies that collect or process data from connected devices to maintain comprehensive data inventories.

Blog Post

LLM-assisted Abbreviation Mining for Legacy Systems

This blog post shows the process of mining abbreviations and discovering first concepts a COBOL legacy mainframe codebase is made of with the help of Large Language Models. It uses Python, pandas and Claude 3.5 Sonnet to generate insights that can be gathered from such a simple thing like a list of files.

Blog Post

How To Build a Data Product with Databricks

Blog Post

Creating data products with Terraform on AWS

Have you heard of data mesh? Are you intrigued by its potential but uncertain how to get started building data mesh and data products? If so, this article outlines a potential approach and delves into the key concepts behind it!

Blog Post

Processing medical study data with Data Mesh technologies

Together with our customer CluePoints, we evaluated new technologies, tools and standards for data storage, data processing, data versioning, and data lineage. These might become useful for refactoring their self-serve data platform.

Blog Post

Data Mesh: Decentralized Data Analytics for Software Engineers

Blog Post

Defect Analysis using pandas

Defect Analysis is a classic analysis technique to get insights into how buggy your system might be. In this blog post, we explore how Defect Analysis works and how we can implement it with a standard data analysis tool from Python: pandas.

News

INNOQ launches Data and AI Consulting Services

News

Women+ in Data and AI Summer Festival