1️⃣ LLMs don't understand the sensitivity of data.
Ask it to "analyze customer data", and it will casually combine PII, logs, internal metrics, and test tables in a single query. It has no concept of what it shouldn't touch. This boundary should exist at the architecture level.
2️⃣ Schema exposure is the security surface.
The more raw tables you expose to the GenAI system, the more unpredictable its queries become. Good systems provide curated semantic layers, not the data warehouse itself.
3️⃣ Prompting is not access management.
Writing "don't access sensitive data" in a system prompt is a recommendation, not control. Management should be implemented through access rights, masked views, and query execution gateways.
4️⃣ Observability is more important when working with AI than with humans.
A human executes several queries.
An agent can execute hundreds in minutes.
If you don't track query patterns and cost spikes in near real-time, you won't notice the problem until you receive an incident report.
A common mistake is treating AI as a smart analyst. It's not.
It's a high-speed query generator without judgment, which needs guardrails and a strict execution layer between it and any critical systems.
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
