Customer data for AI use: What should a vendor take and a customer give?
Negotiating how a technology vendor uses customer data brings up not only privacy and data processing concerns but also confidentiality and trade secrets, in a rather abstract way.

Suppose a manufacturer gives an AI-enabled technology provider access to twenty years of equipment-failure records, quality-control reports, internal operating manuals, exception logs, supplier analyses and troubleshooting playbooks. Before using the material to improve its AI system, the provider removes employee names, customer identifiers, equipment serial numbers and other information that might identify particular people or transactions. The resulting dataset may present far fewer privacy concerns, but the system can still learn which failures the manufacturer considers significant, which patterns it has discovered, which interventions have worked, and how it has learned to distinguish an ordinary production problem from the early warning signs of a serious one. Anonymization may have solved one problem without touching another.
A recently publicized proposed transaction involving the purchase of a large airline’s deidentified enterprise data by a major technology company makes the question especially concrete. The information reportedly includes years of internal communications and operational and business records, and the prospective buyer has said the data could be useful for product development and AI model training. The transaction is unusual in scale, but the underlying issue is becoming ordinary. Technology agreements increasingly ask customers to grant vendors rights to use information not simply to provide a service, but also to analyze performance, improve products and develop artificial intelligence capabilities.
The debate is often framed as a privacy question, and it certainly can be one. Personal information raises privacy and data-processing obligations that cannot be ignored. But once an enterprise begins connecting its internal information and institutional knowledge to AI systems, the negotiation also becomes one about confidentiality, proprietary information and potentially trade secrets.
As early as 2012, I wrote on internetcases about IBM’s decision to restrict employee use of Siri out of concern that sensitive company information could be transmitted to and used by a technology provider to improve its service. The technology has changed dramatically since then, but the underlying concern has not. What has changed is the ability of modern systems not merely to receive or retain proprietary information, but to learn from it and potentially make that learning useful elsewhere.
Anonymization does not eliminate competitive value
Consider again the manufacturer. A competitor might derive little value from knowing that a particular machine failed at 3:17 p.m. on a particular Tuesday. It might derive enormous value, however, from knowing that a particular combination of vibration, temperature and maintenance history tends to predict a failure three weeks later. The first is an individual record. The second is knowledge developed from the relationship among many records.
The distinction becomes more important as AI systems ingest increasingly sophisticated forms of enterprise information. Businesses do not simply possess collections of documents and database entries. They develop classifications, relationships, methodologies and rules for understanding those materials, and some of the most commercially valuable information in an organization may exist not in any individual record but in the structure the company has imposed on thousands or millions of records.
One useful concept here is an ontology. In an enterprise system, an ontology identifies the things that exist within a field and the relationships among them. A manufacturer might distinguish among categories of failures, causes, interventions, equipment types and outcomes in ways that reflect years of accumulated experience. Those classifications are not necessarily obvious from the underlying data because the choices about what to classify, what distinctions matter and what relationships deserve attention can themselves embody institutional knowledge. In that sense, a company’s ontology can reveal how the company thinks.
The same can be true of canonical documents. A canonical document is an authoritative or curated expression of what an organization presently understands to be true about a subject. It may contain an approved methodology, a decision rule, accumulated lessons, historical exceptions, preferred approaches or conclusions developed through experience. In an AI-enabled knowledge system, canonical documents can be particularly valuable because they help transform a generic system into one that understands how a particular organization actually operates.
Removing names or customer identifiers from those materials does not necessarily eliminate their value. Indeed, aggregating information across a large number of transactions can sometimes expose patterns and methodologies that are difficult to see at the level of an individual record. That is why anonymization, although important for privacy and data-processing purposes, should not be treated as the complete answer to the customer’s commercial concern. Anonymization can eliminate identity without eliminating competitive value.
The risk of a knowledge leak
Traditional information-security thinking focuses heavily on data leaving the organization. A file is copied, a database is breached, or a confidential document gets into the wrong hands, and the harm arises because another person obtains the information itself. AI introduces a different possibility because a competitor may never receive the customer’s documents, database or internal manuals and still receive some of the benefit of what those materials taught a vendor’s generally available system.
That is what I mean by a knowledge leak. The customer’s raw information does not necessarily leak in the ordinary sense. Instead, knowledge embodied in that information can move into a shared capability that benefits others. The distinction is important because many contractual safeguards were developed around controlling possession, disclosure and retention of information rather than controlling what a system is permitted to learn from it.
Suppose the manufacturer has developed a superior methodology for identifying quality problems earlier than others in its industry. The methodology may be reflected across internal reports, canonical documents, classifications, historical decisions and the relationships encoded in its ontology. If those materials are used to improve a shared AI system, the vendor may eventually provide other manufacturers with a better way to recognize the same problems. The competitor never obtains the source documents or sees the customer’s confidential records, but it can nevertheless receive some of the benefit produced by the customer’s accumulated knowledge.
That possibility demonstrates why traditional concepts of ownership and deletion do not completely answer the AI question. A contract can provide that the customer owns its data and that the vendor must delete the data at termination, and both provisions remain important. But if the vendor has already used the information to improve a shared model or other generally available capability, deleting the source material does not necessarily delete what the system learned from it.
One way to understand the problem is to think about the movement from individual records to the structures used to organize them, from those structures to institutional knowledge, and from institutional knowledge to a learned capability. As the value moves farther from the original record, the traditional question of who “owns the data” becomes less useful by itself. The more important question becomes what the system was permitted to learn and whether that learning must remain within the customer’s relationship with the vendor.
“Training” can be too narrow a word
This is also why a contractual restriction addressing only AI “training” may not do all the work the parties expect. Information can contribute to fine-tuning, evaluation, retrieval systems, embeddings, benchmarking, feedback processes, synthetic data, prompt optimization, algorithmic improvement and other techniques that may not always be described internally as model training. New methods will appear, and the terminology used to describe them will continue to evolve.
Contract drafting should therefore focus on permitted and prohibited outcomes as well as particular technical processes. If the concern is that confidential customer knowledge should not become part of a capability made available to competitors, the agreement should address that result directly instead of depending entirely on whether an engineering team characterizes a particular process as “training.” This is why it is not enough to negotiate only over who owns the data. The parties should also negotiate over where the learning is allowed to go.
The vendor has legitimate interests
A sensible approach should not begin from the premise that a vendor is wrong to want information from its customers. Technology providers need information to operate their services, maintain security, troubleshoot problems, measure performance and determine whether features work. Product telemetry can help identify failures and improve reliability, while aggregate usage information can provide legitimate insights into how a service should evolve. Customers themselves benefit from many of these activities.
AI can make the vendor’s legitimate interests even stronger. A customer may specifically want a system to learn its terminology, understand its documents, recognize its classifications and improve based on its interactions. Customer-specific retrieval, configuration or model improvement may be central to the value of the service being purchased. A customer that insists that nothing generated through its use of the service can ever contribute to improvement may therefore be asking for something commercially unrealistic or technologically undesirable.
I have written previously on internetcases about the tendency of both sides to over-negotiate technology and AI agreements when the better approach is to identify the particular risks involved and allocate them with greater precision. In that discussion, the concern surrounding customer information was already apparent, including questions about where information is stored, how long it is retained, whether it is used for model training and whether it can be absorbed into broader datasets controlled by a vendor or an underlying model provider. Data rights illustrate the same point especially well because the sensible choice is rarely between giving the vendor everything and giving it nothing.
The important question is what the vendor actually needs to accomplish the legitimate purpose at issue. Operational telemetry may be enough to diagnose whether a feature is working. A customer’s proprietary methodology may not be necessary for that task. Customer-specific retrieval may require access to confidential documents, but it does not necessarily follow that the same documents must contribute to a model made available across the vendor’s customer base.
Stress test “we cannot segregate it”
The question becomes particularly important when a vendor responds to a proposed restriction by saying that its system cannot accommodate it. Perhaps the vendor says it cannot exclude one customer’s information from a shared model, cannot separate material once it enters a training environment, or cannot distinguish substantive customer content from information used to improve the service. Those statements may be accurate, but they should ordinarily begin the next stage of the discussion rather than end it.
The first task is to determine what cannot means. Something may be technically impossible, or it may simply be unsupported by the vendor’s current architecture. A control may be possible but available only in another product tier, or implementing it may require a segregated environment that costs more to operate. In still other circumstances, the asserted technical limitation may actually reflect a business model built around obtaining broad rights to use customer information. Those are materially different situations and should not be collapsed into the same answer.
Customer counsel should therefore ask how information is classified when it enters the system, whether particular repositories can be made unavailable for generalized improvement, and whether the vendor can distinguish operational telemetry from substantive customer content. The inquiry should also address whether proprietary ontologies or canonical documents can remain customer-specific while other information is used more broadly, whether customer information can support retrieval without entering a shared training process, whether customer-specific improvements can remain customer-specific, and what derived representations actually move from the customer’s environment into systems or processes shared with others.
None of these questions assumes that the vendor must be able to satisfy every request. A vendor may have legitimate technical reasons why a restriction is difficult, and segregated infrastructure, customer-specific models or specialized controls may increase development and operating costs. The commercial result may be reduced functionality, a different product tier or a higher price. That is a real negotiation about technology and economics rather than a negotiation cut short by an unexplored assertion of technical necessity.
The larger point is that architecture and contractual entitlement are different things. A vendor can reasonably explain that its chosen architecture has consequences for price or functionality, and a customer can reasonably decide whether those consequences are acceptable. But a vendor’s chosen architecture does not itself determine the scope of rights the customer must contractually surrender.
A more useful negotiating framework
The parties can usually make more progress by breaking the issue into a series of practical inquiries rather than arguing about “AI rights” in the abstract. The process should begin by identifying the information involved. Raw transaction records should not automatically be treated the same as internal playbooks, confidential communications, proprietary taxonomies, ontologies, canonical documents or analytical methodologies. Different categories of information can carry different kinds of value and present different risks. As I discussed recently in connection with structured data licensing, contractual restrictions on access and use can themselves become an important part of protecting valuable datasets and the interests embodied in them.
The next inquiry concerns what the vendor proposes to do with the information. Operating the service, securing the service, troubleshooting problems, improving functionality for the particular customer, generating aggregate analytics, improving a shared model and developing unrelated products are different activities even if a contract attempts to gather all of them under a phrase such as “improving the services.” The parties should determine which activities actually require substantive customer information and which can be accomplished with more limited data.
They should then ask who receives the benefit of the resulting learning. There is a substantial difference between making a system better for the customer whose information produced the improvement and using that information to make the system better for everyone in the market. That distinction may provide one of the most useful dividing lines in the negotiation because it directs attention away from terminology and toward the commercial consequence the customer actually cares about.
Technical feasibility comes next. If the vendor says information cannot be separated, the parties should understand why and what alternatives exist. Where meaningful segregation imposes real costs, those costs can become part of the commercial discussion rather than being concealed inside an overly broad data license.
Only after those questions have been answered should the parties assign the corresponding rights. A customer may readily permit processing necessary to provide the service, protect security and improve functionality within its own environment. Carefully structured aggregate analytics may also be acceptable. The same customer may reasonably object to allowing substantive confidential materials to improve generally available models, methodologies or features that will be made available to competitors.
The vendor should perform the same exercise from the other direction. It should ask whether it really needs the underlying customer content or whether telemetry would suffice, whether the resulting learning needs to become part of a shared system, whether a right must survive forever, and whether every category of customer information must be included. Once the parties separate these questions, what initially appears to be a philosophical disagreement about artificial intelligence often becomes a series of ordinary commercial decisions.
Not everything valuable must be a trade secret
Trade secret law provides an important part of the legal background. Under federal law, qualifying business, technical and other information can constitute a trade secret when the owner has taken reasonable measures to maintain its secrecy and the information derives independent economic value from not being generally known or readily ascertainable by another person who could obtain economic value from its disclosure or use. The statutory definition expressly contemplates information such as compilations, methods, techniques, processes and procedures.
But a technology contract does not need to turn every negotiation into a prediction about whether a particular body of information would ultimately prevail in a trade-secret lawsuit. A customer may possess commercially sensitive information that it reasonably would not provide to competitors even if counsel could not confidently establish every element necessary for statutory trade-secret protection. Contractual confidentiality and use restrictions can protect legitimate business interests broader than the minimum necessary to win a misappropriation claim.
That distinction matters when dealing with ontologies, methodologies and canonical documents. The practical question is not simply whether each item can be labeled a trade secret as a matter of law. It is whether the customer has developed information, structure or institutional knowledge that gives it a commercial advantage and that it would not willingly make available to competitors. If so, the parties should consciously decide whether the vendor needs a right to convert that value into a generally available capability.
Vendors can respect that concern without giving up every ability to improve their technology, just as customers can protect commercially sensitive knowledge without insisting that vendors learn nothing at all. The goal should be precision: the customer should give what the vendor reasonably needs to provide and improve the service, and the vendor should take what it can explain and justify. Both sides should resist language that treats every form of customer information, every AI process and every type of product improvement as though they present the same commercial question.
The most useful question in the negotiation is therefore probably neither “Who owns the data?” nor “Can the vendor train on it?” Those questions remain relevant, but they do not get all the way to the issue that AI has made increasingly important. The better question is: What is the system allowed to learn from this customer, and where is that learning allowed to go?