Showing posts with label COSO. Show all posts
Showing posts with label COSO. Show all posts

Monday, June 1, 2026

Governing What You Build: Five More Takeaways from the COSO GenAI Framework (Part 2 of 3)

This is the second post in a three-part series breaking down the Committee of Sponsoring Organizations of the Treadway Commission's (COSO) report, Achieving Effective Internal Control Over Generative AI. If you missed Part 1, which covers Shadow AI, GenAI-specific risks, the eight capability types, and the 17 COSO principles, you can find it here. The full report is available at coso.org.


Most organizations have crossed the threshold from experimenting with GenAI to depending on it. Outputs are being reviewed by staff, embedded in workflows, and in some cases informing decisions without anyone asking hard questions about how those outputs were produced or what happens when the underlying model changes. The governance gap from Part 1 is not just open. In many cases, it is being papered over with informal habits and unwritten assumptions.

The COSO report offers a corrective to that drift. The five takeaways in this post address the disciplines that separate organizations that are genuinely governing GenAI from those that are simply using it and hoping for the best.


Takeaway 1: Treat Prompts Like Code


One of the clearest and most practical points in the report, reflected in Principle 5 on accountability, is that prompts should be treated with the same rigor applied to any other controlled configuration. There is a tendency to think of prompting as informal, something closer to a conversation than a system design. That framing is a governance liability.

When you build a prompt, whether for a custom GPT, a Copilot agent, or any other GenAI tool, you are defining an input-to-output process. You are making decisions about what the system will do, how it will behave, and what constraints it will operate under. That is code in every meaningful sense. It should be documented, version-controlled, reviewed before deployment, and subject to change management the same way any other system configuration would be.

The report recommends treating prompts, system prompts, retrieval connectors, and transformation rules as governed configurations with version history, approval workflows, and rollback plans. For organizations that have not yet formalized this, the starting point is simply asking: if this prompt changed tomorrow, would anyone know?


Takeaway 2: The Document's Illustrative Examples Address a Real Problem


One of the practical contributions of the COSO report that deserves recognition is its use of concrete examples throughout the text. The clause extraction assistant, for instance, walks through how a legal team deployed GenAI to identify termination clauses in supplier contracts, what went wrong with lower-quality scanned documents, and how controls were redesigned to address it.

This matters because one of the genuine barriers to AI governance adoption is not skepticism. It is a lack of imagination. Many practitioners understand the principles in the abstract but struggle to see how they apply to the work their teams actually do. The report's examples close that gap. They are not hypothetical edge cases. They are the kinds of workflows that exist in finance, legal, compliance, and operations departments right now, and they illustrate both how GenAI can fail and what a well-designed control looks like in response.

If you are building an internal case for governance investment, the examples in the report are ready-made reference material.


Takeaway 3: An AI Policy Is Not Just a Risk Document, It Is a Decision Record



Principle 6 of the COSO framework calls for organizations to specify suitable objectives for each GenAI use case, and one of the most important applications of that principle is the AI policy itself. The report makes the case that having a documented AI policy is not simply about restricting what employees can do. It is about recording conscious, intentional decisions about how the organization has chosen to approach GenAI.

Consider a marketing department that has assessed its GenAI use as low risk and decided to allow broad access. That may be a perfectly reasonable conclusion. But if it is not documented, it is not a decision. It is an oversight waiting to become a problem. A policy that captures the reasoning, the scope, the acceptable use boundaries, and the data classification rules for a given context transforms an informal practice into a governed one. It also provides a baseline for future risk reassessment as the technology and the regulatory environment continue to evolve.

The organizations that will be best positioned as GenAI regulation matures are the ones that can demonstrate not just that they are using the technology, but that they made deliberate choices about how and why.


Takeaway 4: Model Drift Is a Real Risk


Principle 9 of the framework addresses the need to identify and analyze significant change, and in the GenAI context, few risks illustrate this more clearly than model drift. Model drift refers to the gradual or sudden degradation of a model's performance over time, which can result from changes in the underlying data, shifts in the operating environment, or updates pushed by the model vendor.

The practical implication is straightforward and easy to overlook. A prompt that produced reliable, accurate outputs on one version of a model may produce materially different outputs on a newer version, even if the prompt itself has not changed. This is not a theoretical concern. Organizations that have built workflows around specific model behavior need to treat vendor model updates the same way they would treat any other significant system change: test the outputs, document the comparison, and confirm that the behavior you are relying on has been preserved before the new version goes into production.

This requires building model version awareness into your governance process, something most organizations have not yet done. The COSO report's emphasis on continuous risk assessment rather than annual reviews reflects exactly this reality.


Takeaway 5: The Human in the Loop Is Not a Backup Plan. It Is the Control.


Principle 10 of the COSO framework addresses the selection and development of control activities, and the most important of these in the GenAI context is human review. The report is clear that GenAI outputs should be treated as assertions requiring evidence, not facts to accept by default, and that the level of human corroboration should be proportionate to the risk involved.

A useful frame for this is to think of GenAI as a junior employee. You would not take a first-year analyst's work product and send it directly to a client or use it to support a material business decision without reviewing it first. Not because the analyst lacks potential, but because the stakes of an undetected error are too high and the track record is not yet established. GenAI warrants exactly the same posture. The outputs can be valuable and the productivity gains are real, but the human reviewer is not a formality. They are the control.

The report identifies several approaches to operationalizing this, ranging from full re-performance of AI outputs in high-risk scenarios to risk-based sampling in lower-stakes contexts. The right level of review depends on the use case. The wrong answer is no review at all.

In the post, we will conclude the review of COSO's GenAI framework with our final five takeaways.

Reference

Emett, S., Eulerich, M., Guthrie, J., Pikoos, J., & Wood, D. A. (2026). Achieving effective internal control over generative AI (GenAI). Committee of Sponsoring Organizations of the Treadway Commission. https://www.coso.org/generative-ai

Tuesday, May 12, 2026

The Governance Gap Is Already Open: What the New COSO GenAI Framework Tells Us (Part 1 of 3)

This is the first in a three-part series breaking down the Committee of Sponsoring Organizations of the Treadway Commission's (COSO) newly released report, Achieving Effective Internal Control Over Generative AI. Each post covers five key takeaways from the document. Part 1 lays the foundation: the risks, the capability types, and the control principles organizations need to understand before anything else. The full report is available free of charge at coso.org and is worth reading in full. What follows is a guided tour of the highlights.


Generative AI is not waiting for your governance team. It is already inside your organization, running inside productivity tools, shaping analyses, and generating content, regardless of whether your policies have caught up. The question is no longer whether your employees are using it. The question is whether you know how, where, and with what data.

The COSO report opens with that precise tension. It acknowledges the productivity gains and the analytical possibilities that GenAI introduces across finance, compliance, and operations. It also makes clear that those same qualities, speed, accessibility, and adaptability, are exactly what make GenAI a governance problem if left unmanaged. Hallucinations, prompt injection, model drift, opaque reasoning, and rapid configuration changes can all threaten the reliability of operations and reporting if no one is watching.

That framing sets the stakes. And if your organization has not begun building the internal controls to match, the gap between where you are and where you need to be is already widening.


Takeaway 1: Shadow AI Is the New BYOD


History does not repeat itself, but it certainly rhymes. In the early 2010s, the rise of the iPhone and Android forced IT departments to grapple with the Bring Your Own Device (BYOD) movement. Workers wanted their personal devices connected to corporate systems, and IT had to build frameworks to accommodate that demand without compromising security. BYOD ultimately displaced BlackBerry's enterprise dominance because the pressure from the workforce was impossible to contain.

The same dynamic is playing out now with AI, and the COSO report names it directly. On page five, the document defines Shadow AI as unauthorized or ungoverned AI implementations operating outside formal IT oversight.

The parallel to BYOD is instructive, but Shadow AI carries a higher risk profile. Getting corporate data onto a personal device in the BYOD era required some degree of technical sophistication. With Shadow AI, the barrier is copy and paste. An employee can move sensitive client data, unreleased financial projections, or regulated personal information into a consumer AI tool in seconds, without any technical skill and without any visible footprint in your systems.

What makes this particularly hard to contain is that the motivation is legitimate. GenAI tools offer genuine productivity advantages, competitive edge in knowledge work, and time savings that employees feel immediately. That is not bad behavior. It is rational behavior in the absence of a governed alternative. The COSO report is right to surface this in the introduction, because until organizations provide a sanctioned path, employees will build their own.


Takeaway 2: Seven GenAI-Specific Risks


Before the document maps controls to any framework, it lists the risks that make GenAI governance categorically different from traditional IT risk management. These are not generic technology risks. They are specific to how GenAI systems work and how they fail.

The report identifies seven:

  1. Data quality, source, and completeness
  2. Reliability and consistency
  3. Explainability and transparency
  4. Security and privacy
  5. Bias and fairness
  6. Third-party and vendor risk
  7. Governance and accountability

Each of these deserves its own treatment, and later posts in this series will go deeper. For now, the important point is the list itself. These risks are not hypothetical. They are active in any organization where GenAI is being used, whether governed or not. Shadow AI, by definition, means these risks exist without the controls designed to manage them.


Takeaway 3: Eight Capability Types That Map How GenAI Works


One of the most practically useful contributions in the COSO report is its capability-first taxonomy. Rather than organizing GenAI by vendor or product name, which would be outdated before the ink dried, the report organizes it by what the system actually does. This is the right approach. It gives practitioners a durable lens for risk assessment and control design that does not depend on which tools are in the market this quarter.

The report identifies eight capability types following a data-to-decision sequence (Emett et al., 2026, p. 7):

  1. Data extraction and ingestion
  2. Data transformation and integration
  3. Automated transaction processing and reconciliation
  4. Workflow orchestration and autonomous task execution
  5. Judgment, forecasting, and insight generation
  6. AI-powered monitoring and continuous review
  7. Knowledge retrieval and summarization
  8. Human-AI collaboration

A few of these are worth highlighting from a practical standpoint. Data transformation and integration is one of the most powerful and underappreciated capabilities. The ability to take unstructured information and convert it into structured outputs, or take raw data and convert it into a readable memo, is something GenAI does unusually well. This is not simple summarization. It is a genuine transformation of information across formats and registers that previously required significant human effort. I refer to this as "Data to Documentation" within my GenAI workshops. 

Knowledge retrieval and summarization is another that has real-world traction right now. Tools like NotebookLM are already being used to synthesize large document sets into accessible summaries, a task that once took days. The capability is real, and the productivity gain is real, which is exactly why the governance question cannot wait.

Judgment, forecasting, and insight generation is the most nuanced of the eight. It sits at the intersection of classic machine learning and generative AI, and the report acknowledges that complexity. This capability will receive more attention in Parts 2 and 3 of this series, particularly around how the COSO framework addresses the risk of over-reliance and how human review requirements scale with the materiality of the decision.


Takeaway 4: Five Foundational Characteristics That Impact Control Design


Before mapping any of the 17 COSO principles to GenAI, the report establishes five foundational characteristics of the technology itself. These are not risk categories. They are architectural realities that should inform how controls are built. The report's treatment of each is worth reading in full; the short version is below (Emett et al., 2026, p. 8):
  • Probabilistic, not deterministic: GenAI can be confidently wrong; outputs require validation
  • Dynamic: models, prompts, and data change frequently, sometimes without notice
  • Easily scalable: automation scales errors just as readily as it scales quality
  • Low barrier to entry: accessibility is what enables Shadow AI to flourish
  • GenAI can help govern GenAI: its pattern-recognition capabilities can strengthen monitoring and validation

Takeaway 5: The 17 COSO Principles as They Apply to GenAI


The COSO Internal Control Integrated Framework organizes its guidance around five components and 17 principles. The report applies all 17 to the GenAI context. Here is how they break out across the five components (Emett et al., 2026, pp. 5, 9–17):

Control Environment

  • Principle 1: Demonstrate commitment to integrity and ethical values
  • Principle 2: Exercise oversight responsibility
  • Principle 3: Establish structure, authority, and responsibility
  • Principle 4: Demonstrate commitment to competence
  • Principle 5: Enforce accountability

Risk Assessment

  • Principle 6: Specify suitable objectives
  • Principle 7: Identify and analyze risk
  • Principle 8: Assess fraud risk
  • Principle 9: Identify and analyze significant change

Control Activities

  • Principle 10: Select and develop control activities
  • Principle 11: Select and develop general controls over technology
  • Principle 12: Deploy through policies and procedures

Information and Communication

  • Principle 13: Use relevant information
  • Principle 14: Communicate internally
  • Principle 15: Communicate externally

Monitoring Activities

  • Principle 16: Conduct ongoing and/or separate evaluations
  • Principle 17: Evaluate and communicate deficiencies

What the report does that previous frameworks have not is apply each of these principles specifically to the GenAI context, with examples, minimum control expectations, and metrics. A principle like "identify and analyze significant change" reads differently when the change in question is a vendor releasing a model update that silently alters how your automated reconciliation system classifies transactions. The familiar framework is still sound. The terrain it has to cover has changed.

The next two posts in this series continue the conversation, surfacing the report's most relevant guidance for practitioners navigating the governance challenges that GenAI presents.


Reference

Emett, S., Eulerich, M., Guthrie, J., Pikoos, J., & Wood, D. A. (2026). Achieving effective internal control over generative AI (GenAI). Committee of Sponsoring Organizations of the Treadway Commission. https://www.coso.org/generative-ai