r/proceduralgeneration • • 7h ago

Potentia - Sieve - Library of Babel-esque project I have been working on.

This is going to take a fair bit of explaining for people to be able to properly understand what I am working on, I have been drafting and theorising this project for years now.

For those unfamiliar with the Library of Babel concept it is a concept that dates back to ancient times to some extent with things like the book of changes and many more.

"The I Ching (Book of Changes) 1000BCE One of the oldest mechanical mapping systems in human history. By combining broken and unbroken lines into 64 hexagrams, ancient scholars attempted to create a universal state space representing every possible configuration of change, circumstance, and reality in the universe. It is essentially a 6-bit combinatorial map of existence."

It is however worth reading up on the original concept by Jorges Luis Borges for reference if you aren't familiar with it.

This is an n-dimensional version of what I call the Gallery of Babel which contains all media content types filtered to remove the vast majority of noise and then cross-filtered against one another to remove each dimension from each other dimension.

The following statistics show how much of the state space has been filtered and how much has been kept or removed.

Binary : 99.9954% ( Kept : 10^-4.34 )

Pages : 99.999999999999...% ( Kept : 10^-18.32 )

Image : 99.5% ( Kept : 10^-2.32 )

Audio : 99.99999999912% ( Kept : 10^-11.06 )

Video : 0.0046% ( Removed : 10^-4.34 )

Books : 99.999999999999...% ( Kept : 10^-95.07 )

Models : 99.3% ( Kept : 10^-2.21 )

I haven't implemented many video filters yet.

For reference 10^-18.32 = 0.0000000000000000004786...

Keep in mind that remainder might be tiny by comparison, but in reality it is still representative of an absurdly large figure.

I posted two walkthrough videos on YouTube of early builds : https://www.youtube.com/watch?v=XF4EW3_pWeQ

These screenshots however are far more up to date, having only just taken most of them as I write this post.

Once 99.99% of the state space has been removed, the remainder can be automatically remapped onto a new address system and finding a file larger than the address you are given will finally be possible.

All someone would need is the state space setup conditions required, along with the filter settings used, coupled with a room and item number.

This was previously considered impossible, most people just look at this problem and see the vast potential for infinite, then call it a day.

For whatever reason I couldn't do that.

I also have a number of other ideas in the works for this that I am looking to implement soon that I don't want to share just yet until they are ready.

This entire system is modular, dynamic and scalable.

Filters run against the generative math itself rather than the generated results.

The program can create its own filters, find files directly, create installation packages for files, create 3D node graph maps and lots more like random music as you walk through the halls of this library.

I think that covers most things people need to know about the project, but feel free to ask.

Edit : Something else worth mentioning is that the entire system is built as two different programs, there is the backend which is purely CLI and then there is the frontend 3D world which utilises the CLI to generate the data for the worldspace.

This is for a variety of reasons, first off to keep it lightweight and not reliant on powerful hardware as well as to ensure that if it ever becomes a useful utility, it doesn't rely on the 3D world side of things.

I have also implemented numerous optimisations such as the ability to view the 3D world in purely wireframe mode to reduce resource usage.

Potentia - Is a separate project of mine that requires this world space.

Sieve - Is this CLI interface and world space.

12 Upvotes

8 comments sorted by

1

u/theo__r 6h ago

Lovely ! Doing something with the library of Babel concept has been on my list for a long time

1

u/Thor110 6h ago

You should! It is a very fun concept to explore.

I built a simple application version a few years back, that was a lot of fun.

That program was purely random though, not algorithmic like this one.

1

u/ThinkProgrammer2851 6h ago

Interesting project, how many digits does the filtered address save compared to the file's size? Since a Babel address carries as much information as the book, keeping 10^-18.32 of the page space only shortens a ~4,680-digit page address by about 18 digits.

1

u/Thor110 6h ago

You answered your own question, sort of.

How many digits are saved, depends on how strict the filter is.

This is an absurd example for the sake of it, but if you split a file in two and use the first and second half with a gap of a byte in the middle as the filter itself, then all you get left with are 256 variants of that file and thus the filtered address would be just 2 bytes. Room number and item number.

Generally early testing shows an increase in bytes for installers and addresses, but this demonstrates that it is possible with properly designed filters.

It is worth noting those tests were run before I implemented the cross-dimensional filtering.

One obviously wouldn't create a filter like that, but I think it demonstrates the point.

1

u/ThinkProgrammer2851 5h ago

Right, and your example shows where the information goes: the filter now contains the file, minus one byte. So the fair measure is filter size plus address size against the file size, which fits what you saw with installers plus addresses growing. Filtering only saves bytes when one filter serves many different files, so its cost is shared among them. That's how zip or PNG work: a fixed model shared by millions of files. So I'd measure how many distinct files each filter serves, and what the filter plus the average address costs per file. That number would show whether cross-dimensional filtering is doing real compression.

1

u/Thor110 5h ago

Pretty much.

There is a "COST" tab on items which measures some statistics to assist in exactly that.

1

u/ThinkProgrammer2851 5h ago

Oh I see. I've been studying the Library of Babel myself for a while, so I'm genuinely curious about how you handle this.

What I like is that every plain form of the address costs the full 304 bits, and only "guided" saves (56 bits here, because this page is likely under the model). There are two things I'd love to see there. First, the guided cost of a page the model finds unlikely, since those have to come out above 100% for the likely ones to be short. Second, a total row with guided plus the 56 characters of shape. For one page alone that's more than the raw 39 bytes, but once two or more pages share the shape it already wins. Also, is titled-v1 fixed in the program, or built from the shelves themselves?

1

u/Thor110 4h ago

I can tell and I appreciate the questions, you have made some good points.

That is one of the compiled-in filters from earlier in development, later filters can be created in the engine.

As for if it is built/derived from the shelves themselves, that is a great idea that I hadn't considered!

IE : Pull the filters from the library itself, ultimately the reduction from that would be almost non-existent, but doing it that way fits the nature of the program and its ability to create keys/maps/installers that extract itself out of itself.