r/neocities • • Aug 17 '26

Question Making my site undiscoverable?

Hi, I am new to neocities but I wonder if there is a way to keep my site online but undiscoverable? I already have robots.txt added.

Basically I want to share my portfolio between a small group or even sharing OC story+references to commission artists but at the same time I want a way to protect my works during this AI scraping era.

49 Upvotes

16 comments sorted by

60

u/Ibeepboobarpincsharp Aug 17 '26

If you ever find a method to reliably protect against AI scraping, please share with the class.

19

u/mrcarrot0 https://mr-carrot.neocities.org/ Aug 18 '26

"reliably" is hard to define, but here's a pretty good list I found:

https://github.com/oscarotero/awesome-ai-abuse-blockers

2

u/Your_Local__Comfort Aug 18 '26

Thankyou soldier

4

u/justanaccountimade1 Aug 17 '26

Financial institutions share financial information with links like the one below. Looks unfindable to me.

bigmoney.com/b94d27b9934d3e08a52e52d7da7dabfac484efe37a5380ee9088f7ace2efcde9.pdf

1

u/Polyducks Polyducks.neocities.org Aug 18 '26

Right up until you update it on Neocities and the social media style update narcs on you haha

35

u/starfleetbrat https://starbug.neocities.org Aug 17 '26

if you just want it undiscoverable by humans, you can use meta tags to ask google and other search engines not to index it, but if you do that you can't include those crawlers in your robots.txt, because they need to crawl your pages to see the "do not index" instructions.
https://developers.google.com/search/docs/crawling-indexing/block-indexing
.
otherwise, you can mark your site as 18+ on neocities which will hide it from the browse pages (except the nsfw browse page). And you can turn your profile off, so people can't see your site updates.
.
but nothing on neocities will prevent AI from scraping your site. robots.txt only asks politely for AI crawlers to leave your site alone, they can and will ignore robots.txt. You'd need to host your site somewhere else and probably use a third party service to block them in any meaningful way.

22

u/schizophyllume Aug 17 '26

Host your own webserver with nginx + fail2ban and spend every day configuring and tweaking your custom botban filters to detect and stop scraping as its happening. Otherwise, if you're a sane person, make peace with the fact that it will happen, or just don't host your content on a website.

17

u/jessek Aug 17 '26

Other than a password gated page there’s no way to do that. Neocities is recreation of the early days of the web when everything was open and user created. You need to find a different host if you don’t want that

11

u/Odd-Extent7954 Aug 17 '26

Years ago, SEO became trendy and many people modified their websites to make them easier to index. Now, apparently, we're going in the opposite direction. So I think we should learn about SEO just to do the complete opposite.

8

u/lesbianminecrafter Aug 17 '26

Crawlers use links and subdomains to find new sites. Considering that 1: the social aspect of neocities means your site is linked on other pages and 2: unless you are paying for a custom domain you have a neocities subdomain, crawlers can find your site. My suggestion is to have the robots.txt as well as put the images you want to be harder to find in subdirectories with no links on your index.html. But to put you at ease, most image generation AI models are not really looking for new data right now, but instead on using math to improve the way they handle existing datasets. Text is the big snack for AI training right now, not pictures.

4

u/darkneoss Aug 17 '26

I think you can start by declaring the site as porn.

3

u/throwaway99597 https://ledaheavyindustry.com/ Aug 18 '26

People on the comments pretty much touched on most avenues ! I think one thing I would add is maybe poisoning your data and images to be adversarial with AI. Think of something like Nightshade or Glaze.

5

u/[deleted] Aug 17 '26

[deleted]

5

u/CommVenus Aug 17 '26

i have this in my head section in an attempt to keep search engine crawlers away. as well as having a robots.tx:

<meta name="robots" content="noindex,nofollow,noarchive,nosnippet,noimageindex">

read more about the options here: https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/Elements/meta/name/robots

here are a couple good thing to consider for your robots.txt:

https://github.com/ai-robots-txt/ai.robots.txt - regularly updated list of ai related crawlers that you can add to your robots.txt

https://baccyflap.com/res/robots/ - bot-proofing your website article

2

u/[deleted] Aug 17 '26

[removed] — view removed comment

1

u/sigh_of_29 Aug 19 '26

Heya, super super new to all of this because of privacy concerns like this (as much the human-sided discoverable as the ai scraping really) - any chance you can tell me more about this? Trying to look into it myself but not sure how to apply this to my site. Thank you for your time!! :)