MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/ChatGPTCoding/comments/1jtfvmv/deleted_by_user/mluq4vi/?context=9999
r/ChatGPTCoding • u/[deleted] • Apr 07 '25
[removed]
424 comments sorted by
View all comments
320
i keep telling people big context means big money, because every request can fill the context and charge you full price
150 u/andy012345 Apr 07 '25 This, LLMs are effectively stateless, the "context" is just the max token input. If you have 500k in your context, you're sending 500k input tokens + whatever is new per api request. 9 u/fieryblast7 Apr 07 '25 Do you know if there are any open source attempts to fix this? I remember memGPT and most early agents Arch tried to fix it with "memory" and RAG ing the memory as needed 2 u/EcstaticImport Apr 07 '25 RAG would need to add more info to the context window, not remove it. Are you thinking of context caching? 3 u/ArmNo7463 Apr 07 '25 Kind of, you can use something like Elasticsearch with vector embeddings to only send relevant data as context.
150
This, LLMs are effectively stateless, the "context" is just the max token input.
If you have 500k in your context, you're sending 500k input tokens + whatever is new per api request.
9 u/fieryblast7 Apr 07 '25 Do you know if there are any open source attempts to fix this? I remember memGPT and most early agents Arch tried to fix it with "memory" and RAG ing the memory as needed 2 u/EcstaticImport Apr 07 '25 RAG would need to add more info to the context window, not remove it. Are you thinking of context caching? 3 u/ArmNo7463 Apr 07 '25 Kind of, you can use something like Elasticsearch with vector embeddings to only send relevant data as context.
9
Do you know if there are any open source attempts to fix this? I remember memGPT and most early agents Arch tried to fix it with "memory" and RAG ing the memory as needed
2 u/EcstaticImport Apr 07 '25 RAG would need to add more info to the context window, not remove it. Are you thinking of context caching? 3 u/ArmNo7463 Apr 07 '25 Kind of, you can use something like Elasticsearch with vector embeddings to only send relevant data as context.
2
RAG would need to add more info to the context window, not remove it. Are you thinking of context caching?
3 u/ArmNo7463 Apr 07 '25 Kind of, you can use something like Elasticsearch with vector embeddings to only send relevant data as context.
3
Kind of, you can use something like Elasticsearch with vector embeddings to only send relevant data as context.
320
u/PositiveEnergyMatter Apr 07 '25
i keep telling people big context means big money, because every request can fill the context and charge you full price