Artificial IntelligenceCopyright

The enterprise AI legal issue you probably are not thinking about yet

The next portability fight will be about knowledge

aritificial intelligence and intellectual property

Perhaps the strangest intellectual property question to arise in the past few years with AI is not who owns something. It is, rather, whether there is a “something” to own in the first place. That broader question has particular applicability to enterprise adoption of AI technologies.

Imagine that a company uses an AI-enabled, database-driven system for several years. During that time, the AI part of the system does more than just store the company’s documents and data. It classifies information, maps relationships, learns which sources matter, connects terminology used in different parts of the business, and develops increasingly useful representations of how the organization itself works.

Then the company wants to change vendors. The old vendor returns every  byte of the company’s original data. Yet the system now knows something about the company that did not exist when the relationship began. What is that thing? A compilation? A trade secret? An ontology? A set of factual relationships? A machine representation? Is it subject to intellectual property protection? Or merely a capability of somebody else’s software?

That distinction matters because the customer may have gotten all of its data back while still leaving something valuable behind. Before lawyers can decide who owns that something, they may first have to determine what kind of legal object it is, whether intellectual property rights attach to it at all, and whether “ownership” is even the right way to think about the problem.

We spent much of the cloud-computing era worrying about data portability. Enterprise AI may be creating a stranger problem: knowledge portability.

The most valuable asset may no longer be the data

This is not merely philosophical speculation. Deloitte recently described what it calls SaaS tollgating, in which businesses encounter new restrictions or charges when they try to access data held inside vendor-controlled systems for use in their own AI initiatives. Deloitte recounts how a company that had negotiated strong data-access protections years earlier came to regard those provisions as extraordinarily valuable once its system-of-record data became essential to its AI architecture.

That problem concerns getting access to the customer’s existing information. But it points toward a harder one: what happens after an AI system has spent years doing something useful with that information?

There is some history here. As early as 2010, I wrote on internetcases about the emerging importance of data portability in cloud computing, including the need to think about information not only as originally uploaded but also as modified through a service. That problem remains very real. Enterprise AI, however, may be changing the object of the portability inquiry itself.

For some organizations, the most valuable asset may no longer be the data alone. It may also be the structure, context, relationships, and accumulated understanding built around that data. Technology agreements traditionally divide the world into two relatively comfortable buckets: customer data and vendor technology. AI-enabled systems increasingly create something in between.

Call it, for now, the customer-specific knowledge layer. It might include ontologies, taxonomies, knowledge graphs, classifications, relationship mappings, semantic metadata, customer-specific workflows, relevance judgments, embeddings, retrieval indexes, and other structures generated through the interaction between customer information and vendor technology. Some of these things may be subject to intellectual property rights. Others, e.g., simple compilations of facts (think Feist) may not be protectable at all, and still others may embody legitimate rights or interests belonging to both the customer and the vendor.

We have been surprised by ownership before

There is a familiar reason companies might miss this issue. Business intuition about ownership often unknowingly overruns intellectual property law. A company hires an independent contractor to create software, photographs, an illustration, written material, or another deliverable. The company specifies the work, pays the bill, receives exactly what it requested, and quite naturally assumes that it owns what it paid to create.

That assumption can be wrong. Copyright initially belongs to the author, subject to exceptions such as a qualifying work made for hire. For an independent contractor’s commissioned work to qualify under the statutory work-made-for-hire rules, specific requirements must be satisfied. Otherwise, a company that wants copyright ownership generally needs an appropriate transfer of rights.

Enterprise AI may create an even subtler version of that ownership surprise. The customer supplies the data, pays for the system, directs its use, and may contribute enormous amounts of employee judgment and organizational experience. It is easy to assume that whatever valuable thing emerges must belong to the customer. After all, that thing is the collected knowledge or intelligence of the enterprise.

But the analogy breaks down at a revealing point. With the independent contractor, we normally know what the deliverable is. The legal question is who owns the copyright in an identifiable photograph, software program, report, or design. With AI-generated organizational knowledge, we may not yet know what the “deliverable” is, whether it is one thing at all, or whether some of its most valuable elements are even subject to intellectual property rights in the first place.

Companies are already learning to ask who owns AI inputs and outputs, whether their information can be used for model training, and how confidential information will be protected. The question they may not yet be asking is what happens to the valuable layer that emerges cumulatively between the inputs and the outputs.

Copyright starts with a human being

On this issue-spotting quest relating to enterprise AI ownership that we are on, copyright is an obvious realm to explore. But it does not give us an easy answer. One threshold problem is human authorship. In Thaler v. Perlmutter, 130 F.4th 1039 (D.C. Cir. 2025), the D.C. Circuit held that the Copyright Act requires eligible works to be authored in the first instance by a human being.

The U.S. Copyright Office has taken a related approach to generative AI. It has concluded that human-authored expression, sufficiently creative human arrangement, and human modifications may receive copyright protection, while material generated entirely by AI does not become copyrightable merely because a human supplied prompts.

That leaves substantial uncertainty for enterprise knowledge systems. A deliberately human-designed ontology or creatively selected and arranged compilation might contain copyrightable authorship. A machine-generated network of factual relationships might not. Copyright also does not protect ideas, systems, processes, methods of operation, concepts, or facts merely because gathering and organizing them required considerable work. See 17 U.S.C. § 102.

Database law underscores the classification problem. U.S. copyright law may protect sufficiently original selection or arrangement but not the underlying facts, while European law separately recognizes a sui generis database right under specified circumstances. Neither framework maps neatly onto every ontology, embedding space, knowledge graph, or other machine representation an enterprise AI system might generate.

Trade secrets may fit better

Trade secret law may therefore matter more than copyright in many situations. The federal definition is broad enough to encompass patterns, compilations, methods, techniques, processes, procedures, programs, and codes, provided the information derives economic value from secrecy and reasonable measures are taken to protect it. See 18 U.S.C. § 1839.

That is one reason I wrote in 2020 about three ways trade secrets can be more powerful than copyright. Trade secret protection can reach valuable facts, methods, and know-how that copyright does not.

Consider what a mature enterprise AI system might know after years of use. It might reflect which precedents matter to the company, how it categorizes particular risks, how internal terminology corresponds across departments, which facts tend to influence decisions, or which relationships among thousands of records repeatedly prove significant. Individual pieces of information may be unprotectable while the accumulated structure reveals valuable institutional know-how. Simply stated – the company gathers knowledge to do business better.

The complication is that the customer’s know-how may be encoded through the vendor’s proprietary software, models, schemas, ranking methods, and technical architecture. Both sides may therefore have legitimate interests in different aspects of the same functioning system. A provision stating that “Customer owns Customer Data” barely begins the analysis.

“Derived data” may hide the problem

This is where the conceptual problem becomes a contract problem. Technology agreements increasingly use categories such as Usage Data, Derived Data, Service Data, Analytics Data, Feedback, and Improvements, often accompanied by broad vendor rights to use those materials.

Some of those rights are perfectly sensible. A vendor may need telemetry about performance, security, feature usage, and aggregate service trends. But there is a material difference between information about how a customer used the system and information the system developed about how the customer operates.

A record showing that a feature was invoked 4,200 times is usage analytics. A structured representation of how the company evaluates risk, connects information, organizes work, or makes decisions is something different. Sweeping both into a definition of “Derived Data” may conceal the very issue the parties ought to negotiate separately. If the vendor holds onto Derived Data defined this broader way, there is a greater risk of what we have called knowledge leak.

A better framework distinguishes at least five layers: customer source material; customer-specific derived information; customer-specific knowledge structures; machine representations such as embeddings and indexes; and the vendor’s general-purpose technology. There is no obvious reason one ownership rule should govern all five.

Ownership may not be the right objective

This leads to what I think is the most practical insight. A customer usually does not need to own the vendor’s source code, general algorithms, model architecture, or other general-purpose technology. Trying to negotiate ownership of those things may be both unrealistic and unjustified.

The better objective may be reconstructability. If the relationship ends, can the customer obtain enough of its customer-specific knowledge state to create something materially comparable somewhere else?

Depending on the system, that might require export rights for classifications, relationships, ontologies, taxonomies, mappings, metadata, provenance, schemas, configurations, and workflow definitions. It may also require usable APIs, machine-readable formats, documentation, transition assistance, appropriate deletion obligations, and enough continued access to permit an orderly migration.

That suggests a useful question for enterprise AI negotiations. Do not ask only, “Who owns this?” Also ask, “If we leave, what will we lose that we cannot reasonably reconstruct?”

Build versus buy is becoming an intellectual property decision

The same principle should influence system architecture. Enterprises increasingly rely on services such as ChatGPT Enterprise and the OpenAI API, Microsoft 365 Copilot, Google Workspace with Gemini, Salesforce Agentforce, ServiceNow Now Assist, and other AI-enabled platforms that operate against business information and organizational context.

The point is not that these vendors necessarily claim ownership of customer knowledge. OpenAI, for example, currently states that organizations own and control their business data and that business inputs and outputs are not used to train its models by default. The harder architectural question is where the durable representation of what the organization has learned resides.

“Build internally” does not mean every company needs to train a foundation model or operate its own AI infrastructure. An organization can use third-party models for inference while retaining control of its source documents, canonical knowledge, ontology, entity relationships, provenance, classifications, semantic metadata, and other durable customer-specific structures.

An enterprise should be able to rent intelligence tools without making a vendor the only keeper of what the enterprise has learned. That is increasingly both an architecture principle and an intellectual property strategy.

The next portability fight will be about knowledge

There is an interesting irony here. The law is becoming more sophisticated about the old portability problem at almost exactly the moment AI may be creating the next one. The EU Data Act requires providers of certain data-processing services to facilitate switching through measures including open interfaces and machine-readable exports, and switching charges, including data-egress charges associated with switching, are scheduled to disappear on January 12, 2027.

That is meaningful progress. But imagine that moving every underlying record from Provider A to Provider B becomes inexpensive and routine. A company could still arrive at its new provider with all of its files and discover that years of accumulated relationships, classifications, context, and institutional understanding remained behind.

The first generation of cloud contracting taught companies to negotiate over data ownership, access rights, APIs, exports, and termination assistance. Enterprise AI adds two questions that should increasingly appear in contract negotiations and system-design meetings: What new informational assets will this system create from our information and our use? And if we leave, can we preserve or reconstruct them elsewhere?

As enterprise AI matures, those questions may come to seem obvious. Many agreements still try to force the answers into familiar categories such as Customer Data, Derived Data, Feedback, Improvements, and Vendor Technology. The next generation of technology agreements will have to do better, because the enterprise AI intellectual property problem may not begin with deciding who owns “it.” We may first have to decide what “it” is.

 

Need help with a matter?