For site owners
RewearBot
RewearBot/1.0 (+https://rewear.fit/bot)One request at a time to a site, at least 2 seconds apart (longer when your robots.txt sets a Crawl-delay), and at most 20 requests a minute including product photos (up to 60 photos a minute from image servers that hold the photos in a retailer's own product feed).
What RewearBot is
When a network is connected, RewearBot can download product feeds that affiliate networks provide for retailers Rewear works with, and fetch the product photos in them when the retailer’s terms allow.
RewearBot is Rewear's server. Rewear is a closet app: members send it their order emails so the pieces they bought appear in their closet, and a shared catalog of products helps them find and identify those pieces.
RewearBot does two things. It downloads the product photo printed next to an item in an order email a member sent to Rewear, once. And it reads public product pages, the way search engines and shopping sites do, so the catalog has each product's name, brand, category, photos, codes and price.
What it reads from your site
Your robots.txt first, then your sitemaps (the Sitemap lines in robots.txt and /sitemap.xml, including sitemap indexes and gzip files), then product pages listed in them. It looks for schema.org Product data (JSON-LD or microdata) and OpenGraph product tags. When a site has no usable sitemap it may follow a small number of category pages to find product links.
It reads only pages that anyone can see without signing in. It never logs in, fills in forms, adds to a basket, keeps cookies or runs JavaScript, and it never follows links to other sites.
It starts from a product link printed in an order email, from a brand's website when Rewear staff ask for a crawl, or from a site Rewear has approved as a source of product data, which it checks again about once a week. A product page is read again about every 30 days, with If-None-Match and If-Modified-Since so an unchanged page costs you almost nothing.
It keeps product facts only: no personal information. What it reads becomes a public catalog entry only after Rewear staff approve the brand and the site, or the brand that manages its catalog on Rewear accepts it. Brands that manage their catalog on Rewear can turn automatic updates off.
robots.txt and Crawl-delay
RewearBot follows robots.txt as RFC 9309 describes it. It uses the group for "RewearBot" when there is one, otherwise the "*" group, and honours Allow and Disallow with the longest-match rule, "*" and "$".
robots.txt is read again at least every 24 hours. If it is missing (a 4xx answer), RewearBot treats the site as open. If your server fails (a 5xx answer) or does not answer, RewearBot fetches nothing from the site and tries robots.txt again later.
It honours Crawl-delay. A site asking for more than 30 seconds between requests is not crawled at all.
How often it asks
One request at a time to a site, at least 2 seconds apart, or your Crawl-delay when it is longer. Product photos and pages share a limit of 20 requests a minute to one host, and there is a daily limit per site. Image servers that hold the photos listed in a retailer's own product feed may be asked for up to 60 photos a minute.
A 429 or 503 answer makes RewearBot wait for the time in your Retry-After header (up to a day) before asking again. Repeated 403 answers make it stop for the rest of the day. Every request times out after 15 seconds.
How to block it
Add these lines to your robots.txt (shown below). RewearBot will stop reading your pages within 24 hours, the longest it keeps a copy of robots.txt. To block only part of your site, list those paths instead of "/".
RewearBot sends every request from this address: 34.29.50.210 (also listed at https://rewear.fit/bot/ips.json). Anything else claiming to be RewearBot is not us. Please still use robots.txt where you can: it also stops RewearBot from asking.
User-agent: RewearBot
Disallow: /What it fetches from order emails
Only product photos printed inside an item's line of an order email: JPEG, PNG or WebP, at most 8 MB. Each address is fetched once and kept. It removes the address's query string before asking, so tracking parameters are never sent back. It skips logos, icons, tracking pixels and images outside the item's line.
How it identifies itself
User-Agent: RewearBot/1.0 (+https://rewear.fit/bot)
Every request is also signed with Web Bot Auth (HTTP Message Signatures, RFC 9421, Ed25519): a Signature-Agent header pointing at https://rewear.fit, and Signature-Input and Signature headers covering your host. The public keys are at https://rewear.fit/.well-known/http-message-signatures-directory. Bot protection that checks these signatures (Cloudflare, Akamai, DataDome and others) can tell RewearBot from anything pretending to be it, and you can allow or block it by that verified identity.
Contact, removal and opting out
To have a photo removed from Rewear use the takedown form; removed photos are deleted wherever they are used and are never downloaded again. To ask about RewearBot, or to have your site excluded without changing robots.txt, email support@rewear.fit.