KVKK and AI Systems: An Engineering View
Turkish data protection law was written before modern AI. What that means in practice for embeddings, prompt logs, model training and automated decisions.

Turkish personal data protection law, Law No. 6698, establishes obligations that attach to personal data wherever it goes. It was drafted before retrieval systems and large language models were part of ordinary business software, which means it does not name them. It still applies to them.
For teams building AI features into products used in Türkiye, the practical question is not whether the law covers what they are building. It is which components of an AI system hold personal data, and what each one therefore requires.
This is written from an engineering perspective — how the obligations translate into architecture decisions. It is not legal advice, and anyone deploying a system that processes personal data should have the design reviewed by qualified counsel.
The components that hold personal data
The first mistake is scoping the assessment to the database. An AI feature typically creates four or five new places where personal data lives, and several of them are easy to overlook because they are derived rather than stored deliberately.
Source documents. Contracts, correspondence, CVs, customer records, support tickets. This is the obvious one, and it is usually already covered by existing policy.
The vector index. Embeddings are derived from source documents. A vector is not readable text, and that fact leads teams to treat it as anonymised. It is not: embeddings retain enough information about their source that the index should be treated as carrying the same classification as the documents it was built from. If the source contains personal data, the index contains personal data.
Prompt and completion logs. Anything sent to a model can be captured by logging, tracing, caching and observability tooling. Most of these are enabled by default in AI frameworks, and most of them retain content that includes whatever was in the context window. This is frequently the largest unmanaged store of personal data in an AI deployment.
Model training data. If a model is fine-tuned on customer data, personal data has entered the weights. Weights cannot be searched, corrected or selectively deleted. This has consequences discussed below.
Conversation history. Multi-turn systems store prior exchanges to maintain context. That history accumulates personal data over time and rarely has a retention policy attached at the point it is designed.
An assessment that covers only the first of these will describe a system that does not exist.
Legal basis, and why AI changes the question
Processing personal data requires a legal basis. Explicit consent is one, and several others exist, including where processing is necessary for the performance of a contract or for the legitimate interests of the data controller where the data subject's fundamental rights are not harmed.
Two AI-specific problems arise around this.
Purpose limitation and repurposing. Data collected for one purpose cannot simply be reused for another. Customer records collected to deliver a service are not automatically available as training data for a model, and a vector index built to power an internal search tool is not automatically available to power a customer-facing feature. Each new use needs its own basis, and "we already have the data" is not one.
Special categories. Health data, biometric data, and other special categories carry stricter conditions. An AI system that ingests a document set broadly — an email archive, a shared drive — will pull in special category data incidentally unless something prevents it. Ingestion needs filtering and classification at the point of indexing, not after the fact.
Data subject rights meet system design
The rights an individual holds under the law are where AI architecture is tested most directly, because several of them assume the data can be found, corrected or removed.
The right to erasure. A person asks for their data to be deleted. In a conventional system, this is a delete across known tables. In an AI system, it also means removing their content from the vector index, from prompt and conversation logs, from cached responses, and from any backup that will be restored. If the deletion pipeline does not reach all of these, the deletion is incomplete regardless of what the database shows.
The right to correction. Corrected source data has to propagate. An index built before the correction will keep returning the old version until it is rebuilt or the affected entries are updated. Systems that index on a schedule rather than on change will serve stale personal data in the interval.
The right to information. A person can ask what data is held about them and how it is processed. Answering this for an AI system requires knowing which documents contain their data, which indexes were built from those documents, and what the system did with them.
Automated decision-making. The law addresses decisions produced solely by automated processing that produce adverse consequences for the individual. Any AI feature that scores, ranks, screens or classifies people — candidate screening, credit-adjacent assessment, eligibility determination — needs to be examined against this before deployment, not after.
Fine-tuning deserves separate mention here. Data absorbed into model weights cannot be located, corrected or deleted for one individual. A retrained model is the only remedy, and that is a project rather than an operation. This is a strong practical argument for keeping personal data in retrievable stores and out of training sets: retrieval can honour an erasure request, weights cannot.
Cross-border transfer
Personal data leaving Türkiye is governed by its own regime, and the rules were amended in 2024 to introduce mechanisms including standard contractual clauses and binding corporate rules alongside the existing routes, with a notification requirement attached to the standard contract route.
For AI systems this is not an abstract concern, because AI systems call external services constantly. A model API hosted abroad, an embedding endpoint, a reranking service, a hosted vector database, an observability platform, a translation service — each is a transfer if what is sent contains personal data.
Three design responses follow.
Map every external call. Produce a list of every service the system calls, what is sent to each, and where it is processed. This document is the basis for both the transfer assessment and the transparency notice, and it is unglamorous work that no tool produces automatically.
Minimise what leaves. Much of what is sent to external services does not need to contain identifiers. Pseudonymising before the call — replacing names, numbers and identifiers with tokens that are mapped back locally — reduces the exposure without changing the output quality for most tasks.
Use local infrastructure where the requirement is strict. In-country cloud capacity has expanded, which makes it possible to keep storage, indexes and some processing inside Türkiye. That does not resolve residency on its own, because residency is a property of each service in the architecture rather than of the deployment as a whole, but it removes the hardest constraint.
Security obligations
The law requires appropriate technical and administrative measures, and breach notification to the Board and to affected individuals within the prescribed timeframe.
For AI systems, the measures worth naming specifically are these.
Tenant and user scoping inside retrieval. A retrieval call that is not scoped by tenant and by the requesting user's permissions can return another organisation's or another user's personal data, and the model will use whatever it is given. Scoping belongs inside the retrieval call, enforced in code, not applied as a filter afterwards and not stated in a prompt.
Log hygiene. Decide deliberately what prompt and completion logging retains, for how long, and who can read it. The default settings of most AI tooling are wrong for this purpose because they are designed for debugging rather than for data protection.
Access control on indexes. A vector store frequently sits outside the main database's access control regime. It needs its own, matched to the classification of what it holds.
Prompt injection as a security control, not a curiosity. A system that reads documents and then acts will eventually read a document containing instructions. If the system can be induced to retrieve or disclose data outside the requesting user's scope, that is a data breach pathway. Boundaries have to be enforced at the tool and query layer, where content cannot override them.
A practical sequence
For a team adding AI to a product that processes personal data, the order that works:
- Inventory. List every place personal data will exist in the new feature, including the derived stores above.
- Basis. Establish the legal basis for each processing purpose, and check whether the new use is within the purpose for which the data was collected.
- Transfer map. List every external service call and what it carries.
- Rights pipeline. Build erasure and correction so they reach every store in the inventory, and test them before launch rather than after the first request.
- Retention. Set a retention period for logs, conversation history and indexes, and enforce it automatically.
- Transparency. Update the privacy notice to describe the AI processing in terms a reader can understand.
- Review. Have counsel assess the design, and re-assess when the architecture changes.
The step most often skipped is the fourth, and it is the one that produces the worst outcome, because an erasure request that cannot be fulfilled is a visible failure with a regulator attached.
Why this is an engineering problem
Compliance in this area is frequently treated as documentation produced after a system is built. For AI systems that approach does not work, because the obligations — erasure, correction, scoping, residency, retention — are properties the architecture either has or does not.
A system that indexes without recording provenance cannot answer what data it holds about a person. A system that fine-tunes on customer data cannot honour erasure. A system that logs prompts by default has a retention problem it did not choose. None of these are fixable with a policy document.
Designed in from the start, the cost is modest. Retrofitted, each one is a rebuild.
How we approach it
Our platforms are built for organisations operating across borders, which means jurisdiction is a design input rather than a deployment detail. In Clavix360, our AI CRM and ERP system, each organisation is hosted on its own instance with its own database, so scoping and deletion operate on a boundary that is physical rather than logical, and placement can be decided per organisation. Ara, our cross-lingual matching research programme, works against regulatory sources that are public and treats anything a user supplies as a separate class with separate handling.
The general principle we design to: personal data belongs in stores that can be searched, corrected and emptied. Anything that puts it somewhere those operations cannot reach is a decision to be made deliberately, with the consequences understood.

