Recent developments in the Watson Grinding explosion litigation involving 3M have sparked a reevaluation of how organizations govern AI interactions. Central to this case was an engineering expert who utilized ChatGPT to develop his analysis, leading to revealing prompts like “show how 3M is 0% at fault.” Notably, there was no indication that 3M directly instructed the expert on this usage. However, the implications of the AI interactions that arose during the case are significant.
Plaintiffs' attorney Will Moye encountered materials seemingly generated by ChatGPT during his deposition of the expert. After an intense off-the-record exchange, over 350 pages of previously undisclosed ChatGPT content surfaced. This highlighted a critical issue—what originated as a mere analysis suddenly became pivotal evidence in the inquiry.
This scenario resonates with my professional journey, where understanding the nuances behind decisions often transcends the final documents themselves. In complex situations, the real story is frequently found in the minor details: approvals, communication logs, system messages, or process steps that may have been overlooked. With AI now entering the picture, this narrative has expanded significantly.
The Role of Prompts in Decision-making Records
Historically, dialogues surrounding generative AI risks have largely emphasized the importance of controlling inputs. Organizations have advised employees against inputting proprietary code, sensitive data, or any protected information into publicly accessible AI systems. Such measures are indeed foundational, but focusing solely on inputs misses the broader scope of concerns that AI usage encompasses.
Take, for example, an engineer who interacts with an AI system to evaluate different architectural approaches. Throughout this process, they might alter their assumptions repeatedly until the preferred solution emerges. Similarly, a procurement analyst can utilize AI to establish reasoning for a vendor previously chosen. Furthermore, a manager may craft documentation for an employment decision post-factum, while auditors could adjust prompts to downplay deficiencies. These operations do not necessitate malfunctioning AI; the tool can function correctly while still internalizing critical insights, assumptions, and discarded lines of inquiry that never reach the final report.
Nonetheless, it's pivotal to understand that a prompt, when viewed in isolation, doesn't conclusively represent someone's rationale or intentions. Professionals may engage with AI to explore arguments, counter their own biases, or critically evaluate positions they might abandon. It’s this richness of interaction history that can provide necessary context lacking from polished outputs alone.
The American Bar Association has started recognizing the emerging relevance of AI chat histories within discovery processes, noting their potential to capture abandoned theories, inquiries, and narratives that might never feature in a final product. However, beyond discovery, from a governance standpoint, organizations need to clarify when to retain such evidence, who owns it, and the methodologies governing it.
From Input Protection to Lifecycle Management
The 3M case has shifted my perspective on questions surrounding enterprise AI deployments. While it remains essential to scrutinize the information fed into AI models, organizations must also consider the records generated during AI engagement. This line of inquiry moves beyond conventional acceptable-use guidelines into the territory of information lifecycle management.
The lifecycle of an AI interaction doesn't conclude when the output is generated. It's essential to consider the dialogue itself, the results produced, how this material is managed, the duration of its availability, access protocols, and the eventual deletion process. Organizations typically direct their efforts towards initially securing inputs but often neglect the latter stages of this cycle.
Even something as standard as a shared conversation showcases the complexity of these issues. OpenAI’s guidance indicates that anyone with access to shared links can see related conversations. Although search-engine discoverability for shared chats has been removed, the core governance challenge remains: information generated in what appears to be a private space can inadvertently become public.
When you acknowledge that various employees might employ multiple AI tools across the organization—each with distinct retention rules and access controls—the governance complications only escalate. It’s clear that simply implementing basic acceptable-use measures won’t suffice.
While documenting every interaction might seem prudent, my experience tells me that a flood of records does not guarantee effective oversight. Organizations can amass substantial documentation without the capacity to discern its significance or to regulate its usage effectively. Retaining every prompt indefinitely introduces its own suite of privacy and security complications.
A more balanced approach is needed: governance should be influenced by the outcomes of the work. For instance, an employee asking AI to refine an email shouldn't face the same scrutiny as an engineer whose AI-supported analysis could influence safety assessments, or an executive whose decision-making hinges on an AI-generated report. The greater the risk associated with the task, the more scrutiny should be applied to understand AI's role in that process.
For high-impact applications, this suggests retaining sufficient provenance to reconstruct decision-making: relevant prompts, outputs, the AI model utilized, clear records of human review, and context illuminating AI's influence on conclusions. The objective should be to document adequately without inundating the organization with overwhelming amounts of data, thus ensuring clarity when accountability becomes critical.
This approach also necessitates established ownership. Decisions surrounding record management should involve AI teams while also encompassing legal and compliance frameworks—to avoid discovering AI information retention methods only after legal conflicts arise. Chief Information Officers (CIOs) must assume responsibility for delineating what is retained, what is discarded, what can be shared, and how higher-risk AI interactions are integrated into legal and compliance frameworks.
These principles shift focus from mere prompt management to a broader concept of information governance. With AI’s increasing entrenchment in sensitive enterprise functions, understanding the evidence generated by AI interactions is paramount.
Can You Reconstruct the Decision a Year Later?
A valuable habit I've cultivated while navigating operational frameworks is to consider backward from potential future scrutiny. While acknowledging that not every process will falter, posing the question of what evidence would be necessary six months or a year down the line fosters clarity in accountability. If a decision comes under fire, could the organization illustrate what data was accessible, illuminate AI's involvement, identify human responses, and clarify who bore decision-making authority?
The legal perspective on AI interactions is still evolving. It's not guaranteed that all prompts qualify as discoverable evidence; privilege, work-product protection, and contextual nuances can heavily influence material necessity. A New York ruling recently dismissed a request for ChatGPT records based on the contention that such data represented protected legal research. This ongoing uncertainty underscores the necessity for organizations to proactively establish governance policies rather than reacting defensively in the face of scrutiny.
The takeaway from the Watson Grinding case isn't the notion that every AI prompt automatically constitutes a record. Instead, it prompts the question of when an AI interaction merits being treated with the same seriousness as conventional records.
A point made by my colleague, cybersecurity expert Josh Copeland, encapsulates the accountability risks succinctly: “AI won’t testify for you; it won’t take the fall for you; it won’t face penalties; but it can certainly testify against you.”
Ultimately, every organization should have a definitive answer to a critical question: if a decision was scrutinized a year from now, could we recreate the decision-making process that led to it?