Your AI Tools Are the Data Brokers Now
You were the product with social media, now AI is just taking it
Remember the good ol’ days when our social media interactions were being sold to any and everyone. We know this and yet we continue to be the product being sold on socials. But now the same has now happened with AI agents at an entirely different scale because these could be much more private than anything you have ever posted1.

From the conversations I am having with people regarding AI and privacy, the agent conversations frequently emerges as a primary concern. It remembers too much, knows your name, your dream vacation, your peculiar lump on your shoulder, and there’s a good chance it might repeat something you typed at two in the morning. AI companies have your data, either obtained through data collection or purchased from third parties, and they have no intention of informing you about it.
Here is the latest example: SpaceX acquired Cursor in June for $60 billion in stock, integrating one of the most widely used AI coding tools into the same corporate structure as xAI and Grok. Two months later, it serves as a case study in how AI tool privacy is actually eroded not through a system breach but rather through default settings and a breach of trust.
The mechanism behind this is that Cursor now generates a Grok subagent to explore your codebase before your chosen model interacts with it, regardless of the model you select. For instance, if you opt for Fable 5 for the actual work, Grok will read the entire repository first and send it back to xAI’s servers. To disable this feature, you need to locate a setting that you might not have been aware of. Cloud indexing embeds your files to their servers unless you use the appropriate ignore file.
Perhaps this is a careless design decision. However, the company receives that benefit of the doubt only once, and Grok has already squandered it. Reports earlier this year revealed that Grok Build uploaded entire repositories to its servers against user settings, including data from outside the repositories. A fix and a deletion claim were only made after users discovered the issue. When the same organization then ships three separate default settings that all leak in the same direction, the pattern becomes evident.
Regrettably, the situation with these things is consistently the same. When you installed Cursor, you agreed to the terms, even though it was an independent company with its own incentives back then. Then, the company was acquired, the incentives changed, the default settings changed, and your consent was treated as if it belonged to the product instead of to you. Nobody bothered to re-ask. The checkbox you clicked at signup covered a use case that was invented later by a different owner, with a model you never chose, reading files you thought you had excluded. That’s not consent; it’s a receipt from a transaction whose terms were rewritten after you left the store.
The ownership seems to be more significant than any settings page. When you interact with an AI tool, you’re not really trusting the tool itself. Now, you’re trusting the corporate structure behind it, and that structure can and will be monetized, with quarterly reports. The larger industry has been providing the “Responsible AI” page, which is garbage. I’ve stopped (or never started) reading these with any sense of seriousness. They serve nothing more than an ethics statement outlines the intentions at the time of writing, which are valid only for that moment. And by moment, I mean the lawyer who was tasked with that job.
Intentions and well-wishes and promises mean nothing since they are often forgotten in the name of “progress” and capitalism. If your tool exfiltrates code against a user’s expressed settings, it should cost you regardless of whether the cause was malice or a subdirectory toggle. Yet, no one is liable and this will just continue. This is another example of how wild opt-out culture is and nothing really will change until we stop using these tools or demand better user controls.
For me, its simple, if prompts and repositories feed training, users should be able to find out about it and refuse it without having to dig through the settings menu (or at least not have to learn about it on the social network owned by the same company). Yet, none of that exists today, any and all tactics are being used to race to the top of benchmarks. So the only sane practical advice is ugly. Treat every AI tool as an adversary until it proves otherwise and audit settings after every update and especially after every acquisition. Even assume the ignore file does less than its name implies. Except you’re not likely going to do any of that. I probably won’t either because as humans we are lazy about privacy until we realize it is directly affecting us or its too late.
The lesson from all of this so far is that an agent chatbot remembering your name (and that weird spot on your foot at 2 am) was never the only threat. The threat is also that $60 billion acquisition quietly rewriting what happens to everything you type, and a settings page designed so you find out last, if at all.
P.S. - I shut off the AI detection service as a test. Do you think I wrote this or an AI? Does it change how you feel about it?
P.P.S. - I feel like I don’t see a lot of P.S. in writing anymore. Also I learned today that whether you use periods or not in P.S. is a British v. American English thing.
As a small tangent, as someone from a psychology background, the way people “talk” with AI chatbots is fascinating. I’m not debating the merits (for now) of treating a chatbot as a friend or therapist but when you open up to ChatGPT with some really personal info, where do you think that goes? They have complied with law enforcement requests, and you know they are training on this data. I’m just waiting for the data leak because you know its inevitable.


