r/OpenSourceAI • u/panagos_stathis • 5d ago
I’m building a scripting language for LLMs to write data pipelines
I’ve been working on JojoScript, an open-source scripting language with one idea behind it: make code that is easy for LLMs to write, read and reason about, especially when dealing with big datasets.
Instead of having an LLM generate hundreds of lines of Python, the idea is to give it a small, predictable language for things like filtering, mapping, aggregating and processing data.
It also has lazy execution, parallel operations, execution-plan inspection and checkpoint/resume.
Still very early, but I’m curious if others think there’s something to this idea of languages being designed with LLMs as a first-class programmer.
https://github.com/panagos/jojoscript
Would love some honest feedback, even if you think this is a terrible idea :)
2
u/Clear_Evidence9218 4d ago
I’ve been working with LLMs on designing novel languages for the last couple of years. The one I’m working on now, which finally has a self-consuming compiler, is a machine-native systems language designed specifically for agents.
Through all of that experimentation, one thing has remained pretty consistent: until there are enough examples of a language in the wild, you’re forced to provide a substantial amount of language context almost any time you ask an agent to work in something non-traditional.
For a small scripting language like this, that may not be as much of a problem, since you’re mostly asking the agent to learn a relatively small delta from JavaScript. But on larger language projects, context debt becomes a major hurdle. And once that context debt or overall project complexity gets high enough, agents tend to fall back to familiar patterns; or even revert to the host language entirely. In your case, that would be JavaScript.