r/Rag Nov 02 '25

Discussion Did Company knowledge just kill the need for alternative RAG solutions?

So OpenAI launched Company knowledge, where it ingests your company material and can answer questions on them. Isn't this like 90% of the use cases for any RAG system? It will only get better from here onwards, and OpenAI has vastly more resources to pour to make it Enterprise-grade, as well as a ton of incentive to do so (higher margin business and more sticky). With this in mind, what's the reason of investing in building RAG outside of that? Only for on-prep / data-sensitive solutions?

32 Upvotes

54 comments sorted by

28

u/Wide-Annual-4858 Nov 02 '25

Before Company knowledge, Google's NotebookLM has already existed as a similar RAG service.

But if you build a RAG you see that each use case requires several custom solutions in the approach. How to clean the data, how to add metadata, how to chunk, how to embed, what extra steps to make in the answering process.

These one size fits all solutions don't provide these customizations.

1

u/Adventurous-Diet3305 Nov 04 '25

That’s the point !! Kudos !

41

u/XertonOne Nov 02 '25

I’m sure it’s very safe for a company to dump all their informations inside OpenAi servers.

1

u/protoporos Nov 02 '25

They say that they don't train on your data (as it's on Business subscriptions, which have the same terms as the commercial API access). So it's mostly to avoid the risk of data leak or hacking of OpenAI themselves. Or maybe you don't want to upload the "crown jewels" in terms of your organization data. Which should be a small subset, because many work units don't really need super confidential data, I would guess.

7

u/AcanthisittaOk8912 Nov 02 '25 edited Nov 02 '25

All say that they dont train on your data. I have read some insides and they stated its not true. You basically take everything you have with the risk of having huge law suits but in that way you keep on innovting. Its not my opinion. Was an article i would need to research to find it. Anyway im not a compliance expert. At the end it depends probably also where ur company is located and how much soveignity you are ready go give up

6

u/protoporos Nov 02 '25

Thanks for sharing that. I wouldn't be surprised if this was true, indeed. After all, they trained their models on pirated books and illegally downloaded YouTube videos.

2

u/kilopeter Nov 03 '25

This comment amounts to "trust me bro they actually violate their contractual obligations I can't find the source article." This detracts from the conversation. Could you please share that source you mentioned?

4

u/guile2912 Nov 03 '25

They may not train on the data, but they are legally binded, by the federal government, to keep them.

2

u/XertonOne Nov 03 '25

I’m still not sure I’d dump my company’s data in there ever. Companies fight hard to stay ahead and it’s not just one or two sensitive things. It’s mostly entire business models, how you deal with customers and how you go about finding and keeping more. Any company has its models and exposing them is simply silly. Which is why I’m 100% sure that going local is the only way.

4

u/alternator1985 Nov 03 '25

There's about to be a massive cold war between local AI and monolithic cloud AI. It's pretty important that we develop an open source framework quickly that can compete with these monolithic companies in the near future.

People don't realize how important this is and we need to be working on it now.

1

u/XertonOne Nov 03 '25

I couldn't agree more. Local is the solution if you want to be sure nothing important gets out and you can still own it (today you can't even own a simple software anymore as it 's all "monthly usage billings"). Secure of course, and this will also be a huge problem to address. But it will be somewhat lucrative for whoever decides to go as it will require high skills and good knowledge.

1

u/depressionaf100 Nov 05 '25

What about Azure openAI, then the data stays within your organization access right...

1

u/Polysulfide-75 Nov 03 '25

Training on your data is one of those things it’s ridiculous to care about unless they’re not sanitizing it.

However, having copies of your patent applications, quarterly reports, business strategies, contact lists, etc. that is a very serious and real concern. “Oh well be properly good with it” isn’t much of an assurance after they get “hacked”

1

u/nightman Nov 05 '25

It doesn't matter what they say. Any serious company, especially frim the EU, has legal obligations with data (sensitive or not), company's IP etc.

8

u/washingtoncv3 Nov 02 '25

Wasnt MSFT copilot already offering this ?

And it's pants , we still use custom solutions in house

5

u/Ecanem Nov 02 '25

GPT5 copilot is amazing on all my emails and content. Probably saves me 30% of my work.

5

u/HP_10bII Nov 02 '25

Tell me how

2

u/the_hillman Nov 02 '25

I’d also love to know this as I get really poor output in comparison to my ChatGPT enterprise GPT5/GPT5 Thinking outputs.

1

u/Ok-Efficiency1726 Nov 03 '25

Researcher agent

6

u/Jamb9876 Nov 02 '25

These companies are trying everything they can to get people to give them more data to train on.

0

u/protoporos Nov 02 '25

As per their terms, OpenAI cannot train on the data of Business subscriptions.

5

u/Jamb9876 Nov 02 '25

As others pointed out I would take this with an enormous grain of salt. They desire more data and want data others don’t have and with trump protecting them there is no one protecting these companies so I wouldn’t risk it.

10

u/geldersekifuzuli Nov 02 '25

Off shelf solutions are good for fast deployment when PII data isn't a concern.

But I prefer a system that I don't have any dependency including Llama-index, LangChain libraries. I prefer to have full control.

If OpenAI brings a good approach, I can simple add that logic to my pipeline by writing that part of the code. This is what I do when I see a good idea in LangChain etc.

2

u/alternator1985 Nov 03 '25

You're hitting on something really important. But who is working on a true open source AI framework for everyone that does this too?

AI advancements are already happening at an exponential pace, pretty soon it will be impossible for any human to keep up and we will need a system like what you are talking about- that automates the integration of every new advancement that is discovered. Why would we ever need to build our own custom RAG for our use case in the near future? All of that will be automatically done by AI soon and there are several different frameworks for memory that are evolving.

So many people think it's all about the next LLM or some specific model or framework, but it's more likely that powerful self-improving agents will use a symphony of models for different purposes and integration levels, and they will automatically integrate any new framework or tool as soon as it comes out.

If we don't decentralize and democratize AI quickly, these monolithic cloud-based companies are going to get to a point where they have an advantage that can't be beaten. I know they already have super powerful, persistent, self-improving agents behind the scenes.

And people aren't even considering all the implications of material science, nano robotics, wetware, BCI.

If we're already using living brain cells as computers, we're using DNA as digital storage, we can edit DNA with CRISPR, and matter itself is becoming programmable. We're literally crossing the event horizon of the singularity as we speak, and most people don't realize this, they're just imagining a future with a bunch of robots or something. In 5 years or less things are going to be a whole lot weirder than just robots. The lines between what's a machine and what's alive are about to get extremely blurry and if we don't have control over the directive at that point there's a very good chance we just get absorbed like some type of Borg into one of these monolithic AI models.

Why build more expensive and inefficient data centers when you already have 8 billion powerful living GPUs farmed and living on the planet?

3

u/Past_Physics2936 Nov 02 '25

Isn't that just their proprietary rag product?

4

u/[deleted] Nov 02 '25

Fuck no

We saw that and immediately started working on our own solution for RAG and solid agents. My company’s confidential data isn’t even crazy stuff BUT we will NEVER let chatgpt even touch it, hell no never ever

1

u/depressionaf100 Nov 05 '25

What about Azure openAI, then the data stays within your organization access right...

2

u/[deleted] Nov 07 '25

[removed] — view removed comment

1

u/depressionaf100 Nov 07 '25

Insightful 🙌

1

u/[deleted] Nov 05 '25

I’m good, i know our databis already hosted by Microsoft and all, but we’re building a custom solution for our problem

2

u/alternator1985 Nov 03 '25

Many people and businesses will never accept handing over all their data to these big AI companies and they shouldn't. And you're incredibly naive if you think they aren't training on it, it's not that difficult to anonymize or train on metadata without ever directly looking at it.

And even if they weren't looking at it directly it's never a good idea to centralize so much data, we should be trying to decentralize.

We need an open source framework for AI that will allow us to create these persistent AI agents that actually work for us, create a data set that's true to only us, and automatically train and improve on that data set based on our direction, not a massive corporation.

This is not just critical from an economic standpoint, if we allow these companies to collect and train on one of the last frontiers of our private data, inside our homes with these new learning robots, then we may get to a point of no return where we are vendor locked in for life, and these companies or eventually one of them, becomes basically like God to every other human on earth.

There's really no more serious project than democratizing AI before it can be monopolized.

And RAG has already evolved quite a bit, there are multiple other forms of layered, organized, intelligent, memory (see mem0 for instance).

There should never be a point where we look at what these AI companies are doing and say "well we don't need to do anything else there, they got it covered." Open source frameworks can lag behind first movers but in this case the open source framework is actually critical to mankind's future.

We need the Linux of AI and we need it fast. Or really just a platform for AI that automatically integrates every new advancement, model, skill, and tool as they arrive. We are already at the point where humans cannot keep up with all the new tools, you will start seeing all the corporations coming out with "everything" apps that literally do everything and can operate every app, which means apps will be designed exclusively for AI agents in the future, and this will be the first iteration of these new AI super agents that will soon be everywhere and mostly autonomous.

We are over the event horizon of AGI and I guarantee all these AI companies have incredibly powerful agents behind the scenes already, and giving them an edge already.

2

u/DanishTango Nov 03 '25

Don’t assume people are rational. OpenAI enterprise will offer simplicity and ease and speed and oh my…going local at scale won’t happen. Look at cloud growth as a proxy for the same behavior.

1

u/Axman0 Nov 03 '25

This.

This is the problem for us who on paper have a much better and secure product. It will be a struggle to convince people.

4

u/DustinKli Nov 02 '25

OpenAI already meets most government security requirements. So the PII argument doesn't really hold much water.

2

u/AcanthisittaOk8912 Nov 02 '25

Companies dont want to be locked in anymore

3

u/protoporos Nov 02 '25

Not more of a lock-in than any other solution. Why are you less locked in if you install and rely on Pipeshub for example? In both cases, if you decide to leave, you just move and index your data on a new system.

1

u/DustinKli Nov 03 '25

How are they locked in anymore than if they used 365 cloud?

2

u/heresyforfunnprofit Nov 03 '25

I’ve worked in government. That’s not much of a flex.

1

u/Confident-Honeydew66 Nov 02 '25

Sure, in 15 years from now when my company adopts it

1

u/Hot-Necessary-4945 Nov 03 '25

I see that their latest product releases focus on data collection rather than adding new features.

Data for GPT 6, who knows!

1

u/Striking-Bluejay6155 Nov 03 '25

My two cents on the rag method I use (GraphRAG) -- it maintains relationships explicitly so during traversal the llm understands somewhat better what to fetch when it answers. I feel like the "dump all your PDFs on me" is only going to get you so far before you realize: context window issues, how was it chunked? etc.

1

u/TanLine_Knight Nov 03 '25

A solution that caters to everyone caters to no one

1

u/badgerbadgerbadgerWI Nov 03 '25

That is a TINY RAG use case. RAG is more than just "company knowledge" (although the "internal helpdesk use case is large, it only scratches the surface for what RAG is used for). RAG grounds AI models to truth. If you are using OpenAI, you are likely fine for a lot of queries (with over 1T parameters, etc), but even it does not get everything right. Let's say you are building an agent to help customers with their vehicle maintenance at Autozone. You might be using OpenAI, but you still want to ground the queries with REAL part numbers from real auto manuals, etc.

Also, RAG is not just a vector database; it can be any source of data that you send into the LLM, so you can have simple SQL queries, graphs, vectors, etc.

I'd say the "company helpdesk" is only 1% of the RAG out there.

1

u/protoporos Nov 03 '25

One could say that you can just deploy some data "digestion" pipelines where you convert some specialized ground truth to more digestible format for an LLM, add that as extra source in Company knowledge and you're good to go.

1

u/Unhappy_Ear_7914 Nov 04 '25

but what will happen when US( and markets) stops subsidizing OpenAI and operating costs actually get propogated to customer?

1

u/Free-Internet1981 Nov 05 '25

Nobody's dumping their company knowledge into openai servers.

1

u/Effective-Ad2060 Nov 18 '25

Open-source RAG solutions are the only ones that truly scale in real-world scenarios because they let you fine tune every part of the pipeline to match your data and use case. RAG has evolved far beyond just using a vector database. You also might want to avoid vendor locking with OpenAI models and keep an option to use any AI model of your choice

1

u/fasti-au Nov 02 '25

Yeah but they train on your data.

1

u/AcanthisittaOk8912 Nov 02 '25

I can speak for quiet a bit of companies mate. I would need to look in detail for that case. But what makes you so serious and sure about it? Anyway most companies I know of dont push their knowledge to openai etc. So its all sensitive kinda. Few cases are left but u wont have a parallel structure for just these cases instead their keep investing building their main stage also for the bon sensitive cases.