Komprise COO: Clean your data in order to spend fewer tokens

Komprise COO: Clean your data in order to spend fewer tokens

Half of enterprise data is outdated, says Krishna Subramanian — and pointing AI at it is why token bills balloon. The fix starts with the context window

Nicole Deslandes

October 9, 2026    6 Minutes Read


Sometimes an enterprise AI assistant will spend half its budget before its user has typed in a question.

That, says Krishna Subramanian, is what happens when companies wire a model to too many systems and point it at too much data: the context window fills with tool definitions and stale files, and the meter runs on every one of them.

Subramanian has spent a decade watching enterprises lose track of their unstructured data — the documents, video and scans that pile up across file servers and clouds. As cofounder, president and COO of Komprise, she has built a business on finding, classifying and moving it. The proliferation of AI systems in the past few years has turned such housekeeping into more of a priority.

She’s a serial founder: Komprise is the third venture-backed company Subramanian has started with the same two partners, CEO Kumar Goswami and CTO Michael Peercy. She’s also held senior roles at Sun Microsystems and Citrix in between.

Over coffee, Subramanian tells TechInformed why cleaning data cuts token costs, what advice she has for companies still tackling their data and how she uses AI herself.

Where are companies in their AI journeys today?

Most enterprises are still very early, especially when it comes to connecting enterprise data to AI. You have to be careful about the security and governance of that data because it could reveal trade secrets.

Most are dabbling, but interest is growing rapidly. Six months ago, IT people were like, “Well, we don’t know what we’re going to do.” Now they’re like, “We’re creating our own version of an LLM that runs inside our corporation. Now we need to expose our data to it. Now we need to prepare our data.” It’s moved pretty fast.

Why have tokens become such a headache for enterprise AI budgets?

AI is really hard to price because it’s basically helping you with stuff, right? So a lot of AI companies have decided to charge based on how much compute they use. That’s what tokens are. They’re not a direct correlation to users or to value; they’re a correlation to compute, because AI is not software, it’s actually compute. People forget that sometimes.

Maybe tomorrow tokens won’t be the way, and they’ll come up with something else. But somehow they have to charge for the value they’re providing, and the more processing they do, the more they’ll charge you. So the real trick is how you get good answers from AI without processing so much stuff.

For example, a lot of AI connects to your systems through Model Context Protocol (MCP), but if you’re connected to too many things, or the answers you get when you connect are too verbose, the AI has to load all that and keep it in memory. That costs tokens. People are finding a lot of their tokens are spent just loading these MCP APIs. That’s a waste, because you haven’t even asked one question and you’ve used half your tokens.

And unstructured data has billions of files. If you ask AI, “Show me everything in this folder,” and it starts listing a billion files, that’s consuming your tokens too.

What challenges are companies facing with AI?

There’s a lot of experimentation happening. The power and the challenge with generative AI is that it almost feels human. You just talk to it and it gives you answers that sound pretty intelligent. So we think it’s like a human, but it really isn’t. It’s pattern matching across everything we gave it. It’s not actually thinking, even if it looks like it.

It’s also not deterministic. You can ask the same question three times and not get the same answer. And it hallucinates. Sometimes it caves pretty fast. You say, “Wait a minute, is that really correct?” and it says, “Oh, you’re absolutely right.” So there are clearly gaps, hallucinations and misrepresentations, and you need guardrails. Some education is required, but it’s good that people are experimenting and we’re learning what guardrails we need as we go along.

How does organizing data help use fewer tokens?

AI has what’s called a context window, which is kind of like its memory. It loads things there for processing, and it’s finite. For example, if you give some AI a 100-page document, it may only look at the first 10 pages because it doesn’t have enough context window. If you fill it with junk, you get bad results.

With unstructured data, half of the data in any enterprise is outdated. If you point AI at all your data, at least half of it is irrelevant and will contain wrong information. We eliminate that. We give the AI only the most recent, accurate, relevant and curated information, and we feed it in only as the AI needs it. That’s how we reduce token cost: by optimizing the context window.

How do you use AI?

Personally, I use it for a first level of summarizing and finding things. Otherwise I’d have to look at 30 or 40 different sites and aggregate them myself. So maybe I’m using it as a better search.

But I always ask for details and references, and then I read those myself. It’s good at aggregating what’s out there, but I like to look beyond what somebody is saying. They’re saying this, so where are the proof points? Where’s the counterpoint? What does it actually mean? I still have to do that part.

Do AI outputs need to encourage more of that critical thinking, given they’re nondeterministic?

It’s funny how people think AI is only generative AI.

A huge part of AI is machine learning, which isn’t generative and is very predictable and reliable. There’s a lot of value in it for enterprises. We use it in our product for what we call adaptive automation, where the product adapts to what’s happening. That has tremendous value out of the box because it’s predictable.

Generative AI is useful for natural language interaction and in pockets, but for anything of importance, I think you should rely on machine learning. You don’t want different answers to the same question. You want fact-based, accurate responses.

Is there any final advice you’d give companies?

The biggest advice I’d give is to always know your data first. Simply pointing AI at data blind isn’t a good idea in an enterprise. You should classify your data and get some intelligence from it.

Set up AI personas for users that define what data they’re allowed to use, and don’t point AI at raw data. Point it at curated data, because enterprise data is the heart of an enterprise. We haven’t seen it yet, but one mistake here or there could be significant in terms of exposure and liability.

How do you have your coffee?

I make Indian-style coffee. You percolate the coffee overnight, like an overnight drip, and then add it to milk. You can use thin milk. So it’s like a latte.

This interview has been edited for clarity and length.

10 Leaders Defining the Future of Tech

Discover who’s setting the agenda for 2025.

VIEW LEADERS

10 Leaders Defining the Future of Tech

Discover who’s setting the agenda for 2025.

VIEW LEADERS