Saved To My Saved Content

Agentic AI holds the promise of automating the most complex business activities, thus elevating employees into more strategic roles. For that to happen, the work of agents needs to be flawless and reliable.

Yet the reality is that AI agents are hallucinating and introducing errors into workflows, they are declaring activities finished that aren’t, or otherwise failing customer and regulatory scrutiny. Despite agentic AI’s potential, recent surveys indicate that most large companies are not actively scaling agents, and those that are report limited impact on the bottom line.

How can companies create agentic-driven processes that they can rely on, that will scale successfully across the organization, and, most importantly, boost profitability? We believe the answer is to build an operating system, or harness, around agents. This process, called harness engineering, establishes and ensures effective governance, accountability, and proper controls.

As AI models become increasingly commoditized, winning with AI won’t depend on having the best model. It will depend on building an effective harness––one that captures a company’s specific processes, ways of working, and requirements. Few CEOs and their boards are currently building a harness. But they need to act with urgency and make this the focus of their agenda if they hope to lead the agentic AI revolution, rather than fall behind.

How Harness Engineering Helped an Asian Bank

AI agents can deliver results far faster than humans, whether that’s analyzing information, generating code, or providing solutions for customers. But greater efficiency comes with increased risks around reliability and accountability.

By creating an agentic operating system (OS) similar to that of a personal computer, harness engineering mitigates these risks. When your PC runs an application, the OS manages the application so that it doesn’t pose a danger to the overall system. It is the OS that creates user trust. Harness engineering works in a similar way, although human oversight remains a vital element.

The results are impressive Using harness engineering, BCG recently built a fully agentic platform for a large Southeast Asian bank. By doing personalized advisory at scale across human-assisted and digital channels, the platform enabled advisors to triple the amount of time spent on actively engaging with clients. The result was a 30%+ uplift in wealth-advisor revenue productivity, with associated four-to-five-times improvement in customer conversion. In addition, engineering efficiency increased by five times while speed from design to deployment was 50% higher. These enhancements came less from the selected model and far more from the harness around it––the evaluation and approval processes and the behavioral guardrails––which meant the platform could be deployed in a real-world, regulated banking environment.

Using harness engineering, BCG recently built a fully agentic platform for a large Southeast Asian bank. By doing personalized advisory at scale across human-assisted and digital channels, the platform enabled advisors to triple the amount of time spent on actively engaging with clients.

Monthly Newsletter Subscription
Tech + Us: Harness the power of technology and AI

Resetting the Organization

With AI moving at warp speed, organizations often struggle to keep up. In many cases, CEOs are deploying agents at scale but with only limited controls, fragmented governance, and little accountability when things go wrong. To avoid a widespread AI reliability crisis, regulators are starting to move from principle-based guidance for the use of agents to architecture-based requirements.

Regulated industries are a particularly good fit for agents with a harness around them. Such industries as financial services and telecommunications have clear and well-documented standards, making it easier for agents to ensure that companies are acting correctly. However, this requires significant changes in operating models, with humans applying judgment, defining what “good” looks like, and evaluating outputs while orchestrating agents. For their part, the agents produce an audit trail by design rather than as an afterthought, helping to meet compliance obligations. Yet in less regulated settings too, such as software delivery or writing design specifications, harness engineering can also help to create reliability and scale.

By allowing companies to scale agentic AI with confidence, we believe harness engineering marks a major transition in the development of the technology. Up to now, most companies have bolted AI tools, such as copilots and assistants, onto their existing processes. This has increased efficiency at the margins but without significantly altering the degree of human involvement. Using AI agents with harness engineering, companies can reap far greater efficiency benefits.

For example, we were recently employed by a software company to support its transition from an AI-assisted business approach to an agentic-native one. By shifting from using coding assistants to full harness engineering across the product development life cycle––from initial requirements to deployment and operations––we were able to increase the productivity of the team by as much as five times.

Harness engineering involves a fundamental reset in roles, workflows, and processes to create efficiencies. Humans act more as system supervisors: they set goals and boundaries up front, leave the agent to handle routine work, and intervene at critical checkpoints. While such key approval gates still require human involvement, more routine gates are automated, thanks to specialized AI agents that are used to critique worker agents. From our experience in the banking industry, we’ve found that these new structures disrupt traditional workforce pyramids, which contain a large base of junior staff, resulting in a shift toward more demanding design and strategy roles.

Harness engineering involves a fundamental reset in roles, workflows, and processes to create efficiencies. Humans act more as system supervisors: they set goals and boundaries up front, leave the agent to handle routine work, and intervene at critical checkpoints.

Key Parts of a Harness Operating System

There are five key elements that are needed for an effective harness, or agentic operating system.

Specs. The specs are effectively the overarching scope contract that defines the purpose of an AI system and sets the boundaries controlling what agents are allowed to do and how they can achieve their goals. It is machine testable and includes core evaluation criteria. The equivalent in a computer’s OS is the application manifest.

Constitution. The constitution sets the specific rules that govern an agentic system. It determines what actions agents must always take, what actions they must always avoid, what events trigger escalation to a supervisor, and which individual is accountable. The code used to write the constitution is deterministic, ensuring that its requirements are nonnegotiable. The equivalent in a computer’s OS is the kernel.

Control Panel. The control panel is similar to the system monitor or process manager in a computer’s OS. It enables human users to see what the agents did and provides a proper audit trail, ensuring that all agentic actions are compliant with any regulations. It also makes the escalation process more transparent, via an escalation matrix, and includes a failure recovery option in the event of a system error.

Context Hub. Inadequate context is a key reason why agents fail. Without a rich context hub, agents often end up (confidently but incorrectly) guessing when faced with missing information, and they even reach out to external sources via the internet if left unchecked. An effective hub acts as a repository of the background information agents need to understand prompts, and it should contain all the data sources that are relevant to a given project. It is this private context that gives players an edge over rivals. The equivalent in a computer’s OS is shared memory.

An effective hub acts as a repository of the background information agents need to understand prompts, and it should contain all the data sources that are relevant to a given project.

Quality Gates. A common mistake in agentic deployment is treating all quality checks as equal––yet they are not equal. We’ve found that harness engineering works best when there are four distinct types of quality gate: automated gates that prevent programmatic failures; evaluation gates where specialized critic agents evaluate worker agents; human stage gates where a named individual enables the system to mark a given activity as complete; and regulatory gates where a quality agent ensures all outputs meet compliance standards. The equivalent in a computer’s OS is the system call and permissions layer.

In addition to these five components, having a robust governance framework from the outset is a key part of harness engineering. A maturity ladder sets out the clearance rules (which the organization has agreed upon as policy) for different types of content depending on risk level. For example, with sensitive material, such as a client-facing investment proposal, a human may be required to approve every output from the system while battle-tested content or a rough internal draft may involve no explicit human approval. As the outputs delivered by an agentic system become more trusted over time, the type of work involved is typically downgraded on the maturity ladder from “check everything” to “spot check now and then.”

How to Begin

As organizations begin to develop the harness that will enable them to maximize agentic AI’s potential, they should bear the following principles in mind:

Start small. Put simply, a harness is an operating system with multiple interacting components. To ensure these components work effectively with each other and existing processes in a real-world setting, companies need to tread cautiously and test new approaches before scaling up.

Build, don’t buy. Although generic harness platforms are available in the market, we believe a bespoke approach is best. Companies are increasingly buying or renting off-the-shelf AI models. But it is through the harness that a company’s own context, process logic, quality signals, governance requirements, and institutional knowledge start to matter. It is these elements that are unique to an organization––and that competitors cannot simply purchase and replicate. Together with an effective AI model, a customized harness can create a strong competitive moat.

Companies are increasingly buying or renting off-the-shelf AI models. But it is through the harness that a company’s own context, process logic, quality signals, governance requirements, and institutional knowledge start to matter.

Be selective. Not every workflow is ready for agentic AI––or harness engineering. The best early candidates are processes where the work is documented, quality can be checked, and an agentic approach can be applied to existing tools and systems.


Harness engineering is the operating system on which the next decade of agentic-driven work will run, and its importance should not be underestimated. Given the speed with which AI is evolving, CEOs and CIOs must act at pace and make harness engineering, and the organizational changes it demands, the cornerstone of their agentic efforts. That is how companies will tap significant competitive advantage and pull ahead of rivals.