About our news crawler
Transparency for publishers, webmasters, and site operators
Who we are
vancouver.tennis operates a small news-aggregation crawler that fetches public RSS feeds and a limited number of HTML pages to surface tennis news for the Greater Vancouver community. We are not a general-purpose search crawler.
User agent
All requests originated by the crawler include a descriptive User-Agent header:
vancouver.tennis-bot/1.0Requests include an explicit contact address in the same header when possible, and the crawler identifies itself in the From: header as well.
What we fetch
- Public RSS and Atom feeds from a small allow-list of tennis publishers and aggregators.
- Where an RSS entry is a redirect (for example, Google News links), a single follow-up fetch to resolve the final canonical URL.
- Occasional Open Graph metadata for image selection, capped at one request per publisher per hour.
We do not crawl behind paywalls, log in to any publisher, scrape full article bodies beyond what a feed already provides, or rehost wire-service images (AP, Reuters, Getty).
robots.txt adherence
The crawler reads /robots.txt on first contact with a new host and caches the decision for 24 hours. We honour User-agent: *, a specific User-agent: vancouver.tennis-newsbot entry if present, and Disallow / Allow rules as defined in RFC 9309. We also respect Crawl-delay values up to 60 seconds.
Rate limits
- Global concurrency is capped at 8 parallel requests, with a per- host limit of 1.
- RSS feeds are polled no more than once every 15 minutes per source, and most sources are polled hourly.
- Conditional requests (
If-None-Match/If-Modified-Since) are used whenever possible to minimize bandwidth.
Copyright and attribution
Articles that are generated from multiple public sources include explicit attribution (publisher, author, original URL, publication date) and link back to the original reporting. We operate under Canadian fair-dealing conventions for news reporting and commentary and do not reproduce article bodies.
Requesting removal
If you are the owner of a source and would like to be removed from the crawler, or if you have questions about how we use your content, email [email protected]. Include the feed URL or domain you control. We will remove the source from our active list within two business days and confirm by reply.
Questions
For other questions, see our about page or our contact form.