AIAgentic AI

    New whitepaper: Behavioral Drift Is the New Attack Surface for AI Agents

    Theo BergqvistGoran Mladenovski
    Theo Bergqvist, Goran MladenovskiSeptember 23, 20265 min read
    New whitepaper: Behavioral Drift Is the New Attack Surface for AI Agents — Turbotic automation strategy article

    Today, we're publishing a new whitepaper: Behavioral Drift Is the New Attack Surface for AI Agents. It maps the four layers where agents drift in production, the threats that follow, and BASS — a framework for keeping agents secure, governed and trusted.

    Today, we're publishing a new whitepaper: Behavioral Drift Is the New Attack Surface for AI Agents. It is the result of months of research into why AI agents behave differently in production than they did in testing — and what that means for security, governance and trust. The full 28-page paper, written by Goran Mladenovski and Theodore Bergqvist, is free to download below.

    Get the white paper

    Fill in your details to unlock the PDF.

    We'll email you occasional research from Turbotic. Unsubscribe anytime.

    A production agent is rarely just a model. It includes system prompts, retrieval, planner loops, memory stores, tool interfaces, action brokers, orchestration logic, safety filters and third-party skills. Real behaviour emerges from the interaction between all of those parts — not from model weights alone. That is exactly why drift can creep in even when nothing in the model has changed, and why the whitepaper treats drift as a security-boundary problem rather than a quality nuisance.

    What is behavioral drift in AI agents?

    Behavioral drift is the gap between how an AI agent behaved when you approved it and how it behaves today. Sometimes the cause is adversarial: a prompt injection, poisoned memory, or untrusted output from a tool the agent calls. Sometimes it is entirely ordinary: a model update from the provider, context accumulating over hundreds of sessions, or a dependency that changed upstream without anyone noticing. The whitepaper defines a four-layer drift model and a threat taxonomy that covers both — and shows why the layers you cannot easily see are the ones that hurt the most.

    The four layers of drift

    Layer What drifts Typical trigger
    Model Outputs, reasoning patterns, refusal behavior Provider updates, fine-tuning
    Prompt System prompts, instructions, templates Edits, versioning mistakes
    Agent Goals, plans, tool use, memory Context accumulation, poisoned memory
    System Integrations, data contracts, dependencies Upstream changes, third-party skills

    Most governance programmes watch the first two layers and stop there. The whitepaper argues that system drift is the most consequential in production: an agent can behave exactly as designed and still cause an incident, because a tool changed its response format or a downstream system shifted underneath it.

    From one untrusted input to a board-level incident

    One of the central contributions of the paper is a seven-stage escalation model. It traces how a single untrusted input — a poisoned document, a malicious tool response, an injected instruction — can escalate step by step into a security, governance, legal or safety incident. The point is not that every drift ends in disaster. The point is that without runtime supervision, you cannot tell which drift is harmless and which is the beginning of an escalation.

    BASS: a framework for enterprise assurance

    The whitepaper proposes BASS — Baseline, Assess, Secure state and supply chain, Supervise execution — as an operational framework for assuring agent behavior across the full lifecycle:

    • Baseline: define behavioral contracts, so there is an explicit, testable definition of what "correct" looks like for each agent
    • Assess: run risk-stratified regression testing, so the depth of assurance matches the risk of the agent's actions
    • Secure state and supply chain: govern memory stores, tool inventories and dependencies, so drift cannot be smuggled in through the supply chain
    • Supervise execution: apply runtime, action-level supervision, so policy is enforced where actions actually happen

    The accompanying reference architecture follows a simple principle: policy at the edge, telemetry at the core.

    Why benchmarks and red-teams are no longer enough

    A benchmark snapshot is evidence about a moment. A pre-deployment red-team is evidence about a configuration. Neither tells you how the agent behaves after a model update, after a month of accumulated context, or after a third-party skill silently changes. Those are precisely the conditions of production, and they are where behavioral drift lives. Evidence of safety has to be continuous, not one-off.

    Who should read it

    The whitepaper is written for CIOs, CISOs, heads of AI and transformation leaders who are moving agents from pilots into day-to-day operations. If you are responsible for approving an agent once and trusting it forever, this paper is an argument for why that model has to change — and a practical framework for what to do instead.

    Frequently Asked Questions

    What is behavioral drift in AI agents?

    Behavioral drift is the gap between how an AI agent behaved when it was approved and how it behaves in production today. It can be caused by adversarial factors such as prompt injection or poisoned memory, and by ordinary factors such as model updates, context accumulation and dependency changes.

    Is behavioral drift a security problem or a performance problem?

    Both, but the security dimension is the one most organisations underestimate. The whitepaper reframes drift as dynamic security-boundary erosion rather than narrow performance regression, which changes what you have to govern and monitor.

    What is the BASS framework?

    BASS stands for Baseline, Assess, Secure state and supply chain, and Supervise execution. It is an operational framework for enterprise assurance of agent behavior, from behavioral contracts and risk-stratified regression testing to runtime, action-level supervision.

    How do I download the whitepaper?

    Fill in the form near the top of this article with your work email and the 28-page PDF unlocks immediately. The download is free, and you can unsubscribe from our research emails at any time.

    Closing thought

    Agents that drift are not a future risk — they are already in production, and the gap between approved behavior and actual behavior is widening as deployments grow. Download the whitepaper, and if you want help putting governance, monitoring and supervision around agents in your organisation, talk to an expert.

    Is your process ready for AI?

    Find out in 2 minutes with our free Automation & Agent Feasibility Check.

    Get started with Turbotic today

    Discover how Turbotic AI can help you scale automation and AI initiatives with full control and visibility.

    Book a demo