In 2030, someone asks about a transaction from 2008. Will an administrator simply click on a documented record, or will they first have to reboot a legacy system that was shut down long ago? This question determines whether your data is an asset or a liability.
Many companies rely on a data lake to collect all their data in one place. But without governance, the data lake turns into a data swamp: full of data that may have been modified, overwritten, or deleted without anyone being able to verify it.
This article explains when a data lake goes awry, why trustworthiness in manufacturing is business-critical, and how audit-proof archiving transforms a data swamp into a verifiable database.
KEY POINTS AT A GLANCE
|
IN A NUTSHELLA data lake is a central repository that stores raw data from all sources without a fixed schema. This makes it flexible, but without governance, it is also unreliable: data is modified, duplicated, or deleted, and its reliability can no longer be verified. In manufacturing, where traceability and audit requirements apply, a data lake alone is not sufficient. What matters is a data repository that fulfills three requirements simultaneously: retention, deletion, and verification. Audit-proof archiving at the database level, as offered by CHRONOS, makes data immutable, readable over the long term, and verifiable in the event of an audit. |
What Is a Data Lake?
A data lake is a central repository that stores large amounts of data from various sources in its raw format without requiring a fixed schema to be defined in advance. Unlike a data warehouse, which structures and standardizes data during ingestion, a data lake follows the “schema-on-read” principle: the structure is determined only when the data is retrieved, not when it is stored.
This approach has clear appeal. A data lake can store everything—from machine data and test results to images and logs—without requiring a decision in advance about how the data will be used later. This flexibility is valuable for big data analytics and machine learning because it can yield insights that no one had planned for at the time of collection.
In manufacturing, the data landscape is growing faster than our ability to control it anyway: process data resides in the MES, inspection results in spreadsheets, and tool logs in legacy systems. The “data lake” promises to bring all of this together in one place. But this is precisely where the problem begins—one that many only realize too late. Our article on process data management in manufacturing discusses how to build a robust data foundation in manufacturing instead.
| Characteristic | Data Lake | Data Warehouse |
|---|---|---|
| Structure | raw; schema defined only upon reading | Structured upon writing |
| Data Types | All, including unstructured | predominantly structured |
| Flexibility | Very high | Limited, but consistent |
| Typical purpose | Big Data, exploratory analysis | Reporting, fixed analyses |
| Risk without governance | Data swamp, loss of trust | Outdated, but traceable |
When does a data lake turn into a data swamp?
A data lake turns into a data swamp as soon as data flows into it without regulations governing its origin, quality, and immutability. The technical term for this is “data swamp,” and it describes the situation perfectly: a murky repository where everything is stored, but no one can tell which data is accurate and which is not.
The core problem isn’t the volume of data. It’s the lack of trustworthiness. In an ungoverned data lake, data records may have been overwritten, duplicated, altered, or deleted without anyone noticing. If someone runs an analysis later, they may be working with data that no longer reflects what actually happened. The storage is full, but its contents are not reliable.
This may be tolerable for exploratory analyses. But as soon as data is to serve as evidence, it becomes critical. An auditor does not ask whether a value is stored somewhere, but whether it has remained unchanged since its creation and who accessed it and when. A mere data lake cannot answer this question. We explore why you shouldn’t rely on a running database alone in our article on why we prefer not to blindly trust a database.
A full data store is not proof. What matters is not whether the data is there, but whether you can prove that it has remained unchanged since its creation.
Why is trustworthiness in manufacturing critical to business?
In many industries, an unreliable data lake is a nuisance. In manufacturing, it’s a liability risk. The reason lies in the specific documentation requirements that manufacturing companies must meet.
Traceability is the first requirement. Suppliers must be able to document which part was installed in which product, along with the corresponding process values. This chain of evidence is only as valuable as the immutability of the underlying data. Our article on traceability in production explains how such a chain is established.
The second requirement is legal retention obligations. According to the CSP white paper on data governance, legal retention periods can extend up to 50 years, and data must remain retrievable, readable, and verifiable throughout this entire period. A data lake, whose formats and content change over the years, cannot guarantee this.
The third requirement is auditability. Standards such as IATF 16949 in the automotive industry or the GoBD in the tax context require robust proof that processes have been controlled and data has been handled properly. Anyone who has to dig this evidence out of a data swamp only when an audit occurs has, in effect, failed to meet this obligation.
The three obligations regarding structured data: retain, delete, and provide evidence
The CSP white paper on data governance sums up the core of the problem in a single formula: Structured data is subject to three obligations simultaneously—all of which relate to the same dataset and can only be fulfilled collectively.
Retention. Commercial law, tax law, industry-specific regulations, and product liability laws prescribe retention periods ranging from 6 to 50 years. According to the white paper, this retention period applies to the content itself, not to the application in which it was generated. Keeping a legacy system alive just to access its data is therefore not a solution, but rather a cost driver.
Deletion. The GDPR requires, under Articles 5 and 17, that personal data not be retained longer than necessary and be deleted upon request. This is not a contradiction to retention but, according to the white paper, an obligation of equal importance with a different focus.
Provide evidence. Both obligations are worthless if compliance cannot be verified. An auditor does not merely ask whether data was deleted or retained, but also when, by whom, and whether the data record was modified in the meantime. It is precisely this ability to provide evidence that is the point at which an ungoverned data lake fails.
The white paper also assigns typical responsibilities to these three obligations: Retention has traditionally been the responsibility of IT through backup and storage; deletion falls under data protection via policies and lists; and verification is the responsibility of the audit department through retrospective reconstruction. As long as these are three separate processes requiring three different tools, every audit becomes a project. Data governance for structured data brings them together within a single dataset.
Data Lake vs. Audit-Compliant Archiving
The difference between a data lake and audit-compliant archiving is not a gradual one, but a fundamental one. It determines whether data is merely available or also verifiable.
A data lake is designed to be dynamic. Data flows into it, is enriched, rewritten, and then extracted again. While this is intentional for analytical purposes, it renders the data unsuitable as evidence because its integrity cannot be guaranteed.
Audit-proof archiving reverses this principle. It stores data in an unalterable form, in an open, long-term-stable format, and logs every access. In doing so, it fulfills the burden of proof, which a data lake fails to meet. Our article on how to properly select audit-proof archiving software discusses the criteria such a solution must meet.
| Feature | Data Lake Without Governance | Audit-compliant archiving |
|---|---|---|
| Modifiability | Data can be modified and deleted | Stored in an unmodifiable format |
| Format | Any, often proprietary | Open, long-term preservation |
| Access logs | Generally not | logged completely |
| Suitability as evidence | low | High, auditable |
| Compliance | Unregulated | GoBD- and OAIS-compliant |
It is important to distinguish this from a backup, which is often confused with it. A backup protects against data loss, but it can be overwritten and is not designed to be immutable. We explain why backup and archiving are two different things in a separate article.
How do you create trustworthy, verifiable data?
The way out of the data quagmire does not lie in more storage, but in governance for structured data. Four principles are crucial here.
Ensure immutability. Data used as evidence must be stored in such a way that subsequent changes are prevented or, at the very least, fully logged. Without this property, no dataset is auditable.
Choose an open, long-term-stable format. Retention periods spanning decades mean that data must remain readable even after the original system has long since been decommissioned. Proprietary formats tied to a specific application pose a risk in this regard. The OAIS standard for long-term archiving addresses precisely this issue.
Integrate retention, deletion, and verification. As long as these three obligations are handled in three separate tools, every audit remains a project. A governance framework that maps all three to a single dataset makes verification the norm rather than a research task.
Replace legacy systems instead of continuing to operate them. Keeping a specialized system alive solely as a data repository ties up resources in licensing, maintenance, and specialized knowledge. If the dataset is archived in an audit-proof manner, the legacy system can be decommissioned. Our article on replacing legacy systems explains how this can be achieved.
What is the cost of an ungoverned data set?
A data swamp isn’t just a quality issue—it’s a tangible cost factor.
About 70 percent of the data in production databases consists of historical records that are no longer needed for operations but continue to incur costs. According to the white paper, if five to ten legacy systems are still running solely as data repositories, the annual costs can reach the seven-figure range.
| Item | Annual Cost Range |
|---|---|
| Legacy business process or ERP module to be phased out but kept in operation | 50,000 to 200,000 euros |
| Mainframe inventory management system at an insurance company | 1 to 5 million euros |
| Search per audit or information request from legacy systems | Days to weeks of effort |
Added to this is the hidden cost associated with compliance. According to the white paper, when a regulatory inquiry is received, the search begins in legacy systems, exported data, and spreadsheets—a research task that can take anywhere from days to weeks. What appeared on paper to be a fulfilled retention obligation thus becomes an expensive project when the time comes. This is precisely where the business case for audit-proof archiving comes in: it reduces ongoing costs and, at the same time, makes providing proof a matter of just a few clicks.
Data Lake and Data Governance with the Manufacturing OS
CSP’s Manufacturing OS integrates process data management, operator guidance, quality assurance, and audit-proof archiving into a single, shared database. The CHRONOS module is crucial for ensuring data trustworthiness: It handles audit-proof long-term archiving and the decommissioning of legacy systems.
CHRONOS identifies inactive data based on rules, converts it into an open, long-term-stable format, and transfers it to storage systems. Archiving is performed in compliance with GoBD and OAIS standards; the data remains unalterable and remains readable in the long term even after the original system has been decommissioned. In this way, CHRONOS precisely fulfills the burden of proof that a pure data lake fails to meet, transforming a data set into verifiable evidence.
The difference from a data lake is thus clear: A data lake collects data; the Manufacturing OS with CHRONOS makes the relevant data immutable, readable in the long term, and verifiable at the push of a button in the event of an audit. Instead of managing a data swamp, a trustworthy database is created that simultaneously preserves, ensures deletable status, and provides verifiable evidence.
Frequently Asked Questions
What is a data lake?
A data lake is a central repository that stores data from various sources in its raw format without requiring a fixed schema to be defined in advance. The structure is determined only when the data is retrieved. This makes it very flexible for big data analytics, but without proper governance, it is also susceptible to quality and trust issues.
How does a data lake differ from a data warehouse?
A data warehouse structures and standardizes data as it is ingested and is suitable for fixed analyses and reporting. A data lake stores data in its raw form and determines its structure only when it is read. The data lake is more flexible, while the data warehouse is more consistent and traceable.
When does a data lake become a data swamp?
A data lake turns into a data swamp as soon as data flows into it without rules governing its origin, quality, and immutability. Then, although everything is stored, no one can tell which data is correct and unaltered. The storage is full, but its contents are no longer reliable.
Why is data in a data lake often untrustworthy?
Because in an ungoverned data lake, data records can be overwritten, duplicated, modified, or deleted without a trace. This is tolerable for exploratory analyses, but it is unsuitable as evidence for audits or liability issues because the integrity of the data cannot be verified.
What distinguishes a data lake from audit-proof archiving?
A data lake is designed to be modifiable, while audit-proof archiving is designed to be immutable. Archiving stores data in an open, long-term-stable format that cannot be altered and logs every access. In doing so, it fulfills the burden of proof, which a data lake fails to meet.
What does GoBD- and OAIS-compliant archiving mean?
GoBD refers to the German principles for the proper maintenance and retention of books and records in electronic form. OAIS is an international standard for long-term archiving. Archiving that complies with both ensures that data remains unalterable, traceable, and readable over long periods of time.
How do you ensure the traceability of production data?
By storing data in an unalterable, open format, logging every access, and consolidating the obligations to retain, delete, and provide proof within a single database. Audit-proof archiving at the database level, as provided by CHRONOS in the CSP Manufacturing OS, meets these requirements.
Is it possible to decommission a legacy system whose data is subject to retention requirements?
Yes, by archiving the legacy system’s data in an audit-proof manner. The retention requirement applies to the content, not to the application. If the content is archived in a way that is unalterable and readable, the legacy system can be decommissioned, saving on licensing, maintenance, and operating costs.
