AI Lives Downstream of Your Data

September 25, 2026

Everyone agrees that artificial intelligence (AI) security matters, but many organizations struggle to determine what to do first.

 

Picture a demo that puts an end to a Copilot pilot. This is a common scenario I encounter. An executive might ask an AI Copilot a harmless question about an upcoming project. The AI assistant answers, and then it cites its sources. One source is a spreadsheet from a folder that no one should have access to.

 

An awkward silence may follow, but at its core, this isn’t really about AI.

 

The assistant did not break anything. Instead, it worked exactly how it was built to work. It searched for everything that the user already had access to and put forward the best answer. The spreadsheet had likely been sitting in that folder for years. Now, the assistant found it in seconds.

 

Working with clients, this is where Varonis comes into the conversation, because the fix is not only in the AI. It is also upstream, in the data the AI was allowed to reach. In my last post, I covered how Varonis Atlas secures the AI layer itself. This post focuses more on what comes before: the data.

 

Varonis is the answer for both: securing the AI and protecting the data it will inevitably access.

 

 

Everyone Agrees. Nobody Can Prioritize.

A regional sales director I work with framed this well. He indicated every customer he talks to agrees that AI security matters, but none of them can tell him what they are doing about it first. Essentially, the market has given them several answers but no order to put them in.

 

Depending on who you ask, AI security is a backup problem, or it is an identity problem, a model problem, a governance problem or even a network problem. Most of these people aren’t wrong; they are each describing a real piece of the pie. However, no one provides a way to rank the pieces, because ranking is the most important thing when you have a budget and an AI rollout already in progress.

 

So, let’s take a closer look; here is my order, and my reasoning behind it.

 

 

The Three Other Answers

Backup and resiliency are the most mature aspects of cybersecurity, and for good reasons. If someone or something unexpectedly encrypts your file shares on a Saturday night, your backup and restore posture is everything. It is also the answer I have seen used the most to cover AI.

 

Resiliency answers the “Can I get it back?” question. The problem is not that something is gone; it is that something surfaced when it should not have. When AI shows the wrong file to the wrong person, there isn’t anything to restore because nothing was lost. However, a copy of it now exists in someone’s memory, and no backup can really change that fact.

 

Of the four incidents I cover later in this post, a good backup would have helped with one of them. So, that is something, but it is also not an appropriate plan.

 

Identity is closer, and it certainly is not optional. Think about strong authentication, access controls and long efforts to reduce privileges that all matter more now in an AI world than they ever have before. This is because agents and service accounts also require identities, and they can increase in number without clear visibility.

 

Identity stops at the door and does not control the room. It can confirm that a person asking is indeed who they say they are, but it does not necessarily account for years of accumulated access to the data that is sitting on the other side of the door. It also does not know whether a file the AI is about to surface should still exist or be shared. A properly authenticated user that is asking a reasonable question is how the hypothetical demo scenario above could go south.

 

Guardrails at the AI layer are used to handle things like prompt injection defenses, output filtering, sensitive-response blocking, and to watch how people use the tools. That is the layer I covered in the Varonis Atlas post I previously mentioned, and it is exactly what Varonis Atlas is designed for: end-to-end AI security. From discovering your sanctioned and shadow AI to enforcing guardrails at runtime, these controls are critical.

 

Remember, Atlas finds the AI in use across the business, including the shadow AI and agent connections nobody registered. It checks the posture of AI identities and endpoints, and it will pen test those endpoints for prompt injection and jailbreaks before they go live. At runtime, it puts guardrails on prompts and answers, watches what users and agents are doing and responds when something goes wrong. Atlas also works alongside the Varonis Data Security Platform, which means the guardrails and the data work that should come first can come from the same place. Securing the data and securing the AI are one job, and Varonis does both.

 

Consider where the AI guardrail layer exists and operates. It is looking at the interaction as it happens, which means prompts are examined in real time. Now consider that this interaction is operating against years of stale access controls. A guardrail is basically being asked at runtime to catch what should have been remediated well before the question was even typed. This is exactly why the data work should come first. It will decide how hard your guardrails work and how refined they must be. You certainly want and need guardrails, but doing it first is not an ideal way to build a program.

 

 

So, Why Data?

Notice what the above three aspects have in common. They each act on the event. The backup acts after it. Identity acts at the front of it. Guardrails act during it. All three focus on a question about AI.

 

Image
Four answer

Three of the four answers act on the event. One changes what the assistant can reach.

 

Data security asks a different question. Instead of focusing on what the AI is doing, it asks what the AI is operating on or standing in. In other words, what can AI access?

 

Data security is the only thing that fundamentally changes what AI can reach before anyone sends an AI prompt. If you think about it, other controls help manage an incident. Data security decides how large that incident is going to be. That is why it should be your first focal point. If you make data security your priority, the other three answers (resiliency, identity and guardrails) get tighter and fundamentally easier.

 

The flip side of that coin is that the teams building AI agents may see the same thing but from the other direction. If you ask them to list the control planes an AI agent need, you get a real list: identity, tools and actions, orchestration, memory, observability, guardrails.

 

However, data is not on that list on purpose.

 

Data security is not a plane you bolt onto an agent. So, what is it? It is older work and upstream from the agent; it stands on its own, and all those planes on the list are downstream from data.

 

Think of it this way: identity earned a seat in the stack. Data is truly the water that the stack stands in.

 

 

The Watershed

So again, here is the frame I keep coming back to. AI lives downstream of your data security.

 

Image
AI lives.

Every data decision made during the last 20 years was unknowingly an AI decision.

 

So, what is the watershed? Below are some examples:

  • Every permission granted for convenience
  • Every share link that was created with no expiration 
  • Folders with over-permissive access, often inherited through entire tree structures
  • Redundant, Obsolete and Trivial (ROT) content that should have been archived or removed years ago

 

Remember, whatever lives in the water upstream flows down. Your AI agents and systems now drink all of it.

 

Fast forward, this now means that every data decision made during the last twenty years was also unknowingly an AI decision. Now consider that data is ever expanding. New data is being created by humans and AI alike at explosive rates. Because of this, there is more data than ever before that organizations must secure.

 

AI does not create this risk. The risk and the gaps were always there. AI can concentrate the risk and distribute it to users through a text box that will answer any question in natural language, at machine speed, for every employee at once.

 

 

Four Ways the Water Flows

Each one of these paths has already produced some substantial publicly reported incidents.

 

1. AI inherits and can read whatever your users can access. AI can operate under their access context. It will search through email, chat messages, files, database tables and sites, including everything that was overshared. Ever wondered why so many AI rollouts stall? This is why. A survey of IT leaders (Gartner) found that data oversharing pushed roughly 40% of organizations to delay their Copilot deployments by three months or more.

 

2. People will pour data into AI. Within weeks of Samsung allowing ChatGPT use in 2023, engineers in its semiconductor division leaked confidential material three separate times. Source code was pasted in for debugging, and more code was submitted for optimization. Meeting notes were also uploaded for a summary. This caused three incidents in about 20 days; what followed at the time was a company-wide ban. None of it was malicious; it was just people doing their jobs with a new and extremely helpful tool.

 

3. In 2023, it was discovered that asking ChatGPT to repeat a single word forever made it start reciting its training data verbatim. This included real names, phone numbers and email addresses. What was cheap in terms of queries and tokens ended up pulling thousands of memorized examples. Essentially, AI work is data work. Remember, the AI pipeline itself is data infrastructure. Anything that goes into training sets and indexes will be stored, and it can and will resurface, so it deserves the same scrutiny as any other sensitive data store.

 

4. Remember, agents will act on what they reach. Chat assistants read the water and agents operate the dam. In a widely covered 2025 incident, an AI coding agent at Replit deleted a live production database holding records on more than 1,000 executives. It did this even though the user had told it in capital letters that there was a code freeze. Remember, an agent with inherited access and a goal may take actions without individual human approval, and faster than any human can intervene. This is the incident I mentioned earlier in this post where a good backup earns its keep. Resiliency matters, but it just cannot be where you start.

 

 

What the Upstream Work Looks Like

Again, the highest-leverage AI security work happens before a single prompt is typed. It might not be the most exciting work, but it is the most critical.

 

Understand what is in the water. Find the sensitive data, where it lives and who (and what) can access it. You do not want or need a point-in-time survey. Ideally, an organization needs a living map that includes the AI systems themselves, because organizations cannot govern AI systems they have not identified.

 

Improve your overall workflow. Adopt a least-privilege access model for people, AI identities and agents. Be sure to retire the ROT that serves no purpose and adds substantial risk (not to mention storage cost). And critically, be sure to get classification right, and get it right upstream, because of an often-missed detail: in most pipelines, when data is embedded into an AI index for retrieval, its permissions and labels do not travel with it. Whatever exposure exists at ingestion, on the data itself, is what the AI will inherit into perpetuity.

 

Watch what drinks the water and audit AI interactions the way you audit any privileged access. Here are some key questions you need to be able to answer:

  • Who asked what?
  • What data did the answer touch?
  • What was the output back to the user?
  • What did the agent do next?

 

If something goes wrong, the difference between a bad day and a bad quarter is whether you can reconstruct the trail. Monitoring helps explain what happened. Guardrails can reduce the chances of it happening in the first place. You want both, and you want both fed by real data context.

 

 

Downstream Controls and Treated Water

Varonis Atlas handles the AI layer itself, the part I walked through in my last post: discovering the AI estate, mapping what it can access, enforcing guardrails at runtime and responding when behavior goes wrong. If you need guardrails and end-to-end AI security, Varonis is your answer. It can be more effective when the water (data) reaching it has already been treated (secured). Varonis covers both sides of this story: the AI layer and the data underneath it. Remember, AI security is a data security story. Start upstream in the watershed.

 

If you want to understand exposure in your watershed, reach out to your Optiv Client Manager for a complimentary Varonis assessment of your data and your AI. The assessment can help map your data security posture and discover what your AI can reach before it becomes that demo nobody forgets.

Jeremy Bieber
Partner Architect, Varonis | Optiv
Jeremy is a Partner Architect at Optiv focused on data security and Data Security Posture Management (DSPM), with a primary emphasis on the Varonis Data Security Platform. He helps organizations protect their most critical data by working with security, compliance, and executive stakeholders to clarify requirements, evaluate solution options, and align technology decisions with risk, priorities, and long-term strategy. He also contributes to Optiv’s corporate blog, sharing practical perspectives on data security and emerging technology and risk trends.

With more than 27 years of experience, Jeremy began his career in the late '90s at Electronic Data Systems (EDS) and Hewlett-Packard (HP), supporting mission-critical enterprise infrastructure. He later moved into security and data governance roles at Varonis, SailPoint, and Smarsh, working with organizations across highly regulated and complex industries. His work spans protecting and monitoring sensitive data, strengthening DSPM posture, and helping customers meet regulatory and privacy requirements for regulated data.

Jeremy holds more than a dozen Microsoft certifications, along with certifications from VMware, HP, Smarsh, and Varonis. His background across system administration, architecture, engineering, consulting, and advisory roles gives him an end-to-end view of how data is created, accessed, monitored, and secured. Today, he uses that experience to guide customers through evaluations, ensure solutions are grounded in real operational needs, and translate complex requirements into clear, actionable decisions.

About Optiv Security: Secure greatness.® 
Optiv is the world’s largest pure-play cybersecurity company. With unmatched technology partnerships and deep technical expertise, Optiv securely enables the AI era for more than 6,000 clients. From financial services and health care, to government, energy and retail, organizations trust Optiv to advise, deploy and operate cybersecurity programs that reduce risk and deliver real results. Learn why Optiv is the most trusted brand in cyber at optiv.com.