Industry Leadership
Strategic Initiatives
CSA's strategic programs driving innovation in AI, cloud, and Zero Trust.
A public-interest 501(c)(3) dedicated to secure and trustworthy AI.




Industry Leadership
Strategic Initiatives
CSA's strategic programs driving innovation in AI, cloud, and Zero Trust.
A public-interest 501(c)(3) dedicated to secure and trustworthy AI.

CSAI FoundationChaptersEventsBlog
New Training Courses: Turn CSA research into practical skills with self-paced Frontier Ready Training.

Building a Generative AI Layer into Legacy Business Applications

Published 10/08/2026

Building a Generative AI Layer into Legacy Business Applications
Written by David Balaban.

Enterprise resource planning systems deployed a decade or more ago remain fully capable of tracking inventory, processing invoices, and managing shipping schedules. What these systems typically cannot do is respond to a request to summarize the prior quarter's supplier delays in natural language. This gap between operational reliability and current AI capability is increasingly common across established enterprises.

The scale of the underlying exposure is significant. A CSA survey of 418 security professionals found that 82% of enterprises already have AI agents running in environments IT had not officially provisioned. Of those organizations, 65% reported an AI agent-related security incident within the past year. A substantial share of that exposure originates in the legacy systems these agents and integrations ultimately connect to.

Replacing a legacy system of this kind is rarely justified on cost or risk grounds, given the multi-year investment typically required to configure and validate it. Organizations more often add an integration layer on top of the existing system instead. Security reviews of AI adoption typically document how that integration was sourced, whether developed in-house or through generative AI integration services, alongside the rest of the AI stack. Implemented with appropriate controls, this approach introduces AI capability without disrupting existing operations.

This article examines how that integration layer is typically built, the controls required to operate it safely, and the points at which such projects most commonly fail.

 

The Legacy System Is Not the Core Problem

The term “legacy” carries a negative connotation that is not always warranted. A mainframe system that has processed payroll reliably for three decades is not inherently defective; it was designed prior to the existence of modern APIs and cloud infrastructure. The primary source of friction is architectural. Legacy systems are built around fixed screens and relational database tables, while generative AI systems are designed to operate on natural language and produce flexible, unstructured output. These two paradigms do not interoperate without an intermediary layer.

Older SAP implementations, AS/400 systems, and custom applications developed in the 2000s share this characteristic. None were designed to supply structured input to a language model. Addressing the interface, rather than replacing the underlying system, is typically sufficient to resolve the practical limitations associated with legacy infrastructure.

 

Picking an Architecture: Sidecar, Wrapper, or Full Integration

Multiple architectural patterns exist for positioning a language model alongside an existing system, and the selected pattern carries significant downstream implications for maintainability and risk.

A sidecar pattern leaves the legacy application unmodified. A separate service is deployed alongside it, communicating through existing APIs, database views, or, in the absence of an API, a screen-reading tool such as UiPath. This service manages prompt construction, invokes a model through a provider such as Azure OpenAI Service, AWS Bedrock, or Anthropic's API, and returns results through the same channel. The underlying legacy system remains unchanged.

A wrapper pattern extends this further by positioning the integration layer between users and the legacy application. All requests, whether human or AI-generated, pass through a single point. Middleware platforms such as MuleSoft or Boomi frequently already occupy this position, making them a practical location for adding a model call.

 

The Three-Layer Architecture

Most production implementations converge on a three-layer structure. The legacy core, at the base, retains the authoritative data. An integration layer sits above it, comprising APIs, message queues, and event-streaming tools such as Apache Kafka. The AI layer sits at the top, consisting of the model itself along with an orchestration framework such as LangChain. Because each layer can be modified independently, this structure simplifies the process of switching model providers as requirements evolve.

 

Data Quality and Retrieval Requirements

The output quality of a language model is directly constrained by the quality of the data it can access, and legacy databases are frequently inconsistent. Column naming conventions established by a developer in 2011 may carry no meaning today. Customer records commonly exist across multiple systems, with inconsistent formatting of the same underlying entity.

Retrieval-augmented generation, or RAG, has become the standard approach to this problem. Retraining a model on internal data is costly and prone to becoming outdated. Organizations instead retrieve relevant records at query time and provide them to the model as context. This typically requires loading data into a vector database such as Pinecone or Weaviate, along with generating and maintaining the associated embeddings. Snowflake and Databricks are commonly used as staging infrastructure prior to this step.

This work is largely infrastructural rather than novel, but it is foundational to reliable output.

 

Guardrails and Risk Controls

The reliability requirements for a production AI integration differ substantially from those of a low-stakes application such as content drafting. A model that occasionally generates inaccurate output presents a materially different risk profile when operating adjacent to an accounts payable system. Hallucination and prompt injection are documented risks rather than theoretical concerns. The same CSA survey referenced above found that AI agent-related incidents already result in data exposure, operational disruption, and measurable financial loss. This risk profile warrants deliberate control design.

Solid implementations build in a few non-negotiables:

  • Human checkpoint: any action involving financial transactions or customer commitments requires review prior to execution.
  • Audit logging: every prompt and response is recorded, since compliance functions will require this record.
  • Input validation: the model is restricted from accessing data outside its defined scope.
  • Rate limits and cost controls: token consumption is monitored and capped to prevent unplanned expenditure.

Establishing these controls is not the most visible aspect of an implementation. It is, however, often the determining factor between a deployment that passes legal and compliance review and one that does not proceed to production.

 

Adoption Examples Across Industries

Several organizations have already implemented this pattern. Klarna developed an AI assistant on OpenAI's models that manages a volume of customer support inquiries equivalent to several hundred human agents. The assistant operates on top of Klarna's existing systems rather than replacing them. Morgan Stanley integrated GPT-4 into research tools used by its financial advisors, using decades of internal reports as the underlying knowledge base. SAP introduced Joule as a copilot layered across its ERP suite. ServiceNow implemented a comparable approach with Now Assist within its workflow platform. Salesforce incorporated Einstein GPT into Agentforce, embedding generative capability directly into a CRM platform already in wide use.

In each case, the underlying system was retained, and AI capability was added as an additional layer.

 

Common Implementation Risks

Several risks recur across implementations of this kind. Latency is one such factor: a model response requiring several seconds represents a significant deviation from the response times expected in a workflow built around instantaneous database queries. Cost is a related consideration, since token-based pricing scales with usage in a manner that differs from traditional flat-fee software licensing.

Versioning introduces additional complexity, since legacy APIs were not designed with frequent modification in mind. Changing model providers can alter the behavior of prompts calibrated to a specific model. A further consideration is trust. Industry research indicates that security stacks have incorporated AI capability more quickly than confidence in that capability has developed. An inadequately reviewed integration is likely to widen, rather than close, this gap. Scope expansion is a further risk, where a narrowly defined pilot for summarizing support tickets can, without clear boundaries, evolve into a significantly larger initiative.

 

A Phased Implementation Approach

Organizations that successfully implement this pattern typically begin with a narrow scope and expand deliberately.

  1. Select a single, low-risk use case, such as a summarization tool rather than a decision-making system.
  2. Document the location of relevant legacy data and establish a method for accessing it without affecting production systems.
  3. Implement a sidecar service initially, without modifying the legacy codebase.
  4. Incorporate logging and a human review step prior to customer-facing deployment.
  5. Evaluate outcomes over a defined period before expanding scope.

Expansion typically follows once the initial use case demonstrates measurable value.

 

Endnote

Legacy systems are not inherently the obstacle they are frequently characterized as. They are typically stable and well understood, and replacing one solely to introduce a conversational interface is rarely justified.

A more effective approach involves adding an integration layer that respects existing infrastructure. This typically takes the form of a sidecar or wrapper that communicates with a model and retrieves relevant context. Human oversight remains necessary for any process involving financial transactions or organizational trust. With appropriate controls and a measured implementation approach, legacy infrastructure functions less as a constraint and more as a stable foundation for further development.

Unlock Cloud Security Insights

Unlock Cloud Security Insights

Choose the CSA newsletters that match your interests:

Subscribe to our newsletter for the latest expert trends and updates