1 of 46

Fighting the Bot InvasionKeeping Your Koha Catalog Online�

Brian Pichman (ByWater Solutions)

Galen Charlton (Equinox Open Library Initiative)

5 September 2025

The Impact of AI on Libraries

2 of 46

The problem

The avalanche of web crawlers and bots has become the number one technical barrier to keeping online resource available.

Nobody is exempt.

Library catalogs (probably) get no special treatment.

3 of 46

What is a crawler?

  • Computer program that fetches content from a website
  • Usually aims to harvest an entire website
  • Purpose can be benign or otherwise

4 of 46

Crawlers are not new

  • Web search engines have been using crawlers for decades
  • Internet Archive regular crawls content to archive it
  • Libraries sometimes crawl too
    • Koha's OAI-PMH client
    • Aspen Discovery's or VuFind's website indexing
  • Libraries have sometimes encouraged crawling their catalogs

5 of 46

Crawlers need to be managed

  • Consume system resources
    • Especially matters for database-backed websites that are intended to be searched
    • Guess what a library catalog is!
  • Can make requests far faster than humans can

6 of 46

Management in the good old days

  • robots.txt
  • Checking "user agents"
  • Blocking addresses

# do some indexing, but don't index search URLs

User-agent: *�Disallow: /cgi-bin/koha/opac-search.pl

7 of 46

The good old days

Public domain images

8 of 46

Today

Gemerated by ChatGPT

9 of 46

"AI"/LLM bots/crawlers

  • Like previous generations of bots, general purpose is to gather content from websites
  • Some to gather web content to train large language models
  • Others may be invoked in real time as users interact with chat assistants

10 of 46

Breakdown of the rough social order

  • A LOT of AI crawlers have cropped up to repeatedly crawl websites
  • And they increasingly don't play by the rules
  • Leading to the worst game of whack-a-mole ever... more on that in a bit

11 of 46

Impact

  • A catalog can start receiving many, many times the traffic it's used to
  • Sometimes functionally indistinguishable from a DOS (denial of service attack)
  • More expense
    • Bandwidth
    • Server resources
    • Mitigation measures
  • Loss of availability

12 of 46

Overall impact to GLAM web resources

  • More expense = less willingness to make GLAM materials openly available
  • Mitigation measures can put barriers in front of legitimate users
  • Increases the difficulty of self-hosting

13 of 46

Mitigations

  • Polite requests
  • Withdrawing the resource
  • Blocking
  • Increasing the cost to crawlers

Public domain image

14 of 46

Polite requests - alas!

"Please, sir, could you not crawl my cataloging a hundred times a minute?"

"Sure... I shall crawl it a thousand times a minute!"

15 of 46

robots.txt does not work

  • Was always advisory, but most of the big crawlers respected it
  • In fact, some of the major "AI" companies do, though some have been caught ignoring it
  • But it's not enough: nowadays the best assumption is that robots.txt will be ignored
    • Doesn't hurt to set it up, though - just keep expectations low

16 of 46

Withdrawing the resource

  • A corporate library's Koha catalog likely doesn't need public access
  • Some academic libraries already require that one authenticates first to use their discovery
  • The Koha system preference OpacPublic can be turned off
  • Of course, taking the catalog off the open net or requiring that users log in to search it isn't practical for the majority of Koha libraries

17 of 46

Blocking

  • When and why to block
  • Network/IP-level blocking
  • Web server/load balancer-level blocking

18 of 46

More precisely, why block?

  • Point is to ignore or drop the bot's request before it gets expensive
  • If all the crawlers wanted is a copy of the Koha logo, a hundred times a minute, wouldn't be a big deal
  • OPAC searching is where it can get expensive for the servers, but any Koha page rendered through Plack + Apache adds up

19 of 46

Network/IP blocking

  • Tells the bot to go away before it makes a request by closing or dropping the connection
  • Can be done at level of a single IP address or ranges
  • Country/region-level blocks are possible
    • IE: Allow only US Traffic
  • Many ways to do it;
    • iptables
    • Network Layer

20 of 46

IP block whack-a-mole

Block a country...

... bot requests start coming from U.S. IP addresses

Block a range or some ASNs...

... bot requests start coming from thousands of U.S. residential IP addresses, one or two requests per address

Block common cloud providers (AWS, Alibaba, Google Cloud)

... oops, you may have just blocked a service you're integrating with

21 of 46

Your adversaries...

  • Are well motivated; they assume that a lot of money can be made via their LLMs
  • Are often well-resourced, including access to botnets
  • Don't care if your catalog in particular has become unusable
  • Might consider libraries' well-structured data to be appetizing

22 of 46

Blocking at the web server/load balancer

  • Idea is to receive the request and evaluate - cheaply! - whether to pass it along for Koha to respond to
  • Can be implemented in various places
    • Apache
    • Forward proxy
    • WAF (Web Application Firewall)
    • Load balancer

23 of 46

Inspecting the request

  • "User Agent"
    • Sent by the web browser to identify itself
      • A malicious actor can set a user agent to whatever they like with ease – in other words they can mimic good traffic if looking at User Agents
    • E.g., "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:142.0) Gecko/20100101 Firefox/142.0"
  • Referrer
  • URL parameters

24 of 46

User Agents

  • Examples of major ones include PerplexityBot/1.0, ChatGPT-User, etc.
  • Sadly, few crawlers are polite enough to send a UA of "EvilBot_Please_Block_Me/666"

25 of 46

Example of blocking by user agent

# Block some bots

RewriteEngine on

RewriteCond %{HTTP_USER_AGENT} "ahrefsbot|amazonbot|anybot|bytespider|claudeb

ot|dataforseobot|dotbot|facebookexternalhit|friendlycrawler|gptbot|meta-external

agent|mj12bot|petalbot|seekportbot|semrush|serpstatbot|yandexbot" [NC]

RewriteCond %{REQUEST_URI} "!403\.pl" [NC]

RewriteRule "^.*" "-" [F]

26 of 46

Blocking inhuman behavior

  • Common pattern for many LLM crawlers is to repeatedly run catalog searches that vary the parameters
  • Sometimes, the requests follow a pattern that you can identify as suspicious
  • E.g., recently a bot that did a bunch of requests that contained "opac-search.pl?count=20"
    • ... which is not a pattern that a real human would do unless they are fiddling with the address bar

27 of 46

Request blocking whack-a-mole

Block a particular user agent

... bot will make them up

Block an old, outdated, out-of-support browser's user agent

... oops, that's what the public access computers you don't have the budget to upgrade are using

Block an inhuman pattern

... eventually, the bot operators will adapt

28 of 46

Increasing the cost to crawlers

  • If you know it's a bot, can have some fun
    • E.g., bot gets its response... after 10 seconds
  • Tarpits like Nepenthes dial that approach up to 11
    • Nepenthes aims to feed lots of nonsense to the bots
  • Tradeoff:
    • Such tactics require that you have the system resources to burn on the purpose
    • Not exactly in the library's job description

29 of 46

How Easy It Is To Scrape Websites For Data

You can point tools (including ChatGPT Like tools) to a website and harvest content and data from it…no permission needed

30 of 46

Tools and strategies

31 of 46

Cloudflare

  • Sits on the edge network
  • Traffic is routed through a variety of data centers that CF owns before reaching the data center– providing “cleaner” traffic
  • Free Tier – great for small basic websites
  • Enterprise Tier- needed for complex sites with a need to allow specific types of traffic, whitelist good traffic patterns, etc.

32 of 46

Anubis

  • Open Source
  • Sits in front of website
  • Increases the cost of crawling by making the user agent solve a "proof of work" problem first
  • Simple to deploy
  • One caveat: its default branding may not be suitable for most libraries; an unbranded version is available

33 of 46

Facing the onslaught

34 of 46

System tuning

  • Small silver lining: the bot invasion gives system administrators an opportunity to look at tuning their Koha system
  • Let's discuss how bots can make things slow

35 of 46

Lifecycle of a request

Load balancer/proxy

Apache worker

Plack worker

36 of 46

Tuning parameters to look at

  • Number of Plack workers
    • You have enabled Plack, right?
  • System memory
    • Point is to avoid high swap utilization, which can really slow things down
    • Influenced by number of Plack workers and Apache workers that are permitted
    • Can be better to let requests drop than overcommit memory

37 of 46

Monitoring

  • System resources (memory, CPU, number of processes)
  • Open connections
  • Number of requests over time
  • Purpose: try to identify a bot invasion ahead of time

38 of 46

Log analysis

  • Proxy (if available) logs
  • Apache logs - e.g., /var/log/apache2/access.log
    • Requests that get to Apache
  • Plack logs - e.g., /var/log/koha/SITE/plack.log
    • Requests that get to Plack
  • Purpose is to identify sources of traffic to see if additional mitigations are needed or if something is slipping through that shouldn't

39 of 46

Catch things like this …

Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives

https://blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives/

40 of 46

Staff interface vs. OPAC performance

  • Lots of ways to set up Koha systems, but many configuration are such that if the OPAC is slow or unavailable, the staff interface is as well
  • The "AI" crawlers generally aren't going after the staff interface as such, fortunately
    • Turning off the System Preference for OpacPublic to Disable
    • If already using Discovery Layer, you can close out the Koha OPAC
  • Separation is possible if necessary

41 of 46

Tactical mitigations

  • While IP and UA blocks are not a long-term solution, in the short term applying them can be useful to restore access to a system
  • Example: analyze logs to see if particular ranges are sending requests, then block them
  • Ditto for URL patterns and user agents

42 of 46

Other ideas

  • Enhance Koha to add support for CAPTCHAs for Koha's search page
    • Problematic, since it can create accessibility barriers
    • But may be a necessary evil
    • E.g., Cloudflare Turnstile, reCAPTCHA
  • Mirrors what's been done in Evergreen and VuFind

43 of 46

To sum up

  • Bots and crawlers are a real problem
  • Mitigations are available
  • Active management is essential: everybody involved is playing either whack-a-mole or cat-and-mouse

44 of 46

Questions?

(and comments!)

45 of 46

Reading list

46 of 46

Thanks!

Galen Charlton�Implementation and IT Manager�Equinox Open Library Initiative�https://equinoxOLI.org/gmc@equinoxOLI.org

Brian Pichman�Head of Systems�ByWater Solutions

bywatersolutions.combrian.pichman@bywatersolutions.com