So it would be effective at preventing your site from being used as training data.
Comment on A Project to Poison LLM Crawlers
eru@mouse.chitanda.moe 4 days ago
i would imagine companies would just filter it out
need some more clever way of hiding it or allow it to be self hosted so that it has various urls
CileTheSane@lemmy.ca 3 days ago
GamingChairModel@lemmy.world 4 days ago
If I am reading this correctly, anyone who wants to use this service can just configure their HTTP server to act as the man in the middle of the request, so that the crawler sees your URL but is retrieving poison fountain content from the poison fountain service.
If so, that means the crawlers wouldn’t be able to filter by URL because the actual handler that responds to the HTTP request doesn’t ever see the canonical URL of the poison fountain.
In other words, the handler is “self hosted” at its own URL while the stream itself comes from the same URL that the crawler never sees.