Lumetrace / AI drift and rogue AI
AI Assurance

When your own AI drifts, or goes rogue.

Two different failures that most vendors sell as one. Rogue is your approved AI doing something it was never approved to do. Drift is your approved AI still doing its job, worse. We report them separately, because the evidence for each is different and pretending otherwise is how honest products stop being honest.

The distinction that decides what anyone can honestly claim

Rogue is a scope question. Did this system talk to a service, a model or a tool that nobody approved? Did it act at three in the morning with no human involved? Did its reach inside your network quietly grow? These are questions about behaviour and destination, and network observation answers them well.

Drift is a quality question. Is the same approved system giving worse answers than it did a month ago? That is a question about content, and content sits inside TLS. A passive sensor can produce strong leading indicators and it can date the moment something changed, but it cannot grade an answer. Any vendor turning packet metadata into a quality score is inventing a number.

So the product does two things instead of one. It reports rogue behaviour as detection, and it reports drift with an explicit statement of how much of the picture is actually available.

Rogue AI: what we detect on the network today

All of the following work from observed traffic, with no agent and nothing in your application path. Most of them key off an approved-AI register that you declare once, and which we seed from what is already happening rather than asking you to write an inventory from memory.

  • Shadow AI. A service genuinely in use that your register never declared. An empty register means no register, never that nothing is approved, so adding this does not turn every service you use into a finding on day one.
  • Zombie AI. A service with a real history that has then gone silent for weeks. The egress path and the credential are very likely still open, and nobody is watching a system nobody remembers.
  • Destinations, models and tools outside the register. Declaring what is approved is what makes the word approved mean anything, and each of the three is checked separately.
  • New external tool servers reached by your agents. An agent acquiring a new capability from outside your control is a scope change whether or not it was malicious.
  • Unattended autonomy. AI activity outside your declared working hours, in your declared timezone. The alert states the basis it used, so you can argue with it.
  • Consumption runaway and loop pathology. Volume that keeps climbing, and agents caught talking in circles. We report consumption rather than currency, because the network does not know your contract price.

Drift: leading indicators, dated, with the confidence stated

  • Provider endpoint and client change. A change in the negotiated TLS characteristics of a provider, a new client fingerprint for a host and service pair, or a host reaching a provider it had never used. These are dated facts, and the finding says explicitly that they do not prove the model changed.
  • A rolling numeric baseline with an honest warm-up. It reports nothing at all until it holds fourteen daily samples, and says so while it is learning. We would rather show you a blank panel for two weeks than a confident number derived from three days.
  • Sustained change, not a bad afternoon. A shift has to persist across days before it is reported, and it is expressed as a percentage change with the date it began, because nobody can act on a sigma value.
  • With gateway metadata, certainty replaces inference. If your AI calls pass through a gateway that can export metadata, a genuine model or model version change becomes a stated fact rather than a hypothesis, and token, truncation, refusal and tool-failure rates become measurable.

Providers roll models without announcing it. If you take one thing from this page: the single most valuable field for answering why did it get worse last Tuesday is the model version, and it comes from your gateway, not from the wire.

What we will not claim

This list is deliberately part of the product description rather than something you have to extract from us in a meeting.

  • We do not detect hallucination, prompt injection, insecure output handling or reasoning failure. Those live in the prompt and tool-call layer, which no passive sensor sees.
  • We never call a number derived from packets a quality score.
  • A fingerprint or endpoint change proves the endpoint changed. It does not prove the model changed, and the finding says so in those words.
  • We do not collect your prompt text. Content fields are rejected on arrival by name, not filtered later, and that is a design constraint rather than a policy we could quietly relax.
  • Systems we genuinely cannot see, such as an assistant running in a laptop editor that never crosses a monitored segment, are marked out of scope rather than allowed to pass as clean.
Questions

AI drift and rogue AI: straight answers

Can you really tell me my AI is getting worse?

We can tell you, with dates, that something about it changed: the endpoint it talks to, the client making the call, the shape and volume of the traffic, sustained across days rather than for one afternoon. Whether the answers got worse is a judgement about content, and content is encrypted. If your AI calls pass through a gateway that exports metadata, that gap closes substantially, because a model version change becomes a fact.

What is the difference between rogue and drift, in one sentence each?

Rogue is your approved AI doing something it was never approved to do, which is a scope and security problem. Drift is your approved AI still doing the approved thing, worse, which is a quality problem. The first is detectable from the network; the second needs the network for indicators and your gateway for certainty.

Do we have to declare an inventory of our AI systems first?

No, and we would not trust one written from memory anyway. The register is seeded from the AI activity we already observe, and you confirm or correct a draft. Nothing is rejected for being incomplete; a register with three fields is better than none, and the product reports how complete it is.

Why is the dashboard quiet when we first switch this on?

Because most of it is deliberately silent until it has history. Change detection needs a prior observation, a dormant service needs weeks of silence to qualify, and the baseline needs fourteen daily samples before it asserts anything. That is the design. A product that produced confident AI findings on day one would be making them up.

Find your blind spots before an attacker does.

Tell us whether you suspect something is already inside, or whether you want to measure what your current tools would catch. Either way you get a concrete next step.

Contact us
Related

More on what we detect