{"id":1163,"date":"2012-07-24T15:06:50","date_gmt":"2012-07-24T19:06:50","guid":{"rendered":"http:\/\/michaelnielsen.org\/blog\/?page_id=1163"},"modified":"2012-08-10T13:12:24","modified_gmt":"2012-08-10T17:12:24","slug":"oss-bot","status":"publish","type":"page","link":"https:\/\/michaelnielsen.org\/blog\/oss-bot\/","title":{"rendered":"OSS-bot"},"content":{"rendered":"<p>OSS-bot is a crawler I (<a href=\"http:\/\/michaelnielsen.org\">Michael Nielsen<\/a>) built for educational purposes &#8212; I run occasional informal meetups where programmers in Toronto get together to talk about machine learning, information retrieval, and similar topics.  The crawler:<\/p>\n<p>(1) Is designed to be polite &#8212; it obeys <a href=\"http:\/\/www.robotstxt.org\/\">robots.txt<\/a>, as well as various other best practices.  If you wish to exclude OSS-bot, please add the appropriate exclusion rules to your robots.txt file, for some User-agent pattern matching the string &#8220;OSS-bot&#8221;.  Note that the crawler caches robots.txt files, so it may take up to a day or so to cease accessing your site.  Contact me (mn@michaelnielsen.org) if this is a concern.<\/p>\n<p>(2) The mean time between requests is usually about 3 minutes, for any given domain.  The crawler will never take less than 70 seconds between requests for any given domain.<\/p>\n<p>(3) Crawls typically last 1-3 hours.  Occasional crawls may last longer, up to 48 hours.  The crawler is run infrequently, and most months is run for between 0 and 10 hours.<\/p>\n<p>(4) Only html content is crawled (not images, javascript, etc), so the total bandwidth consumed is typically a few hundred kilobytes per hour per domain.<\/p>\n<p>OSS-bot crawls in this way so that the burden imposed by the crawler is (briefly) comparable to that imposed by a moderately intense user. However, I&#8217;m still learning best practices for crawling, and don&#8217;t want to impose an undue burden on sites.  If you have concerns or suggestions or would simply like OSS-bot to stop crawling your site, please contact me (mn@michaelnielsen.org).<\/p>\n","protected":false},"excerpt":{"rendered":"<p>OSS-bot is a crawler I (Michael Nielsen) built for educational purposes &#8212; I run occasional informal meetups where programmers in Toronto get together to talk about machine learning, information retrieval, and similar topics. The crawler: (1) Is designed to be polite &#8212; it obeys robots.txt, as well as various other best practices. If you wish&hellip; <a class=\"more-link\" href=\"https:\/\/michaelnielsen.org\/blog\/oss-bot\/\">Continue reading <span class=\"screen-reader-text\">OSS-bot<\/span><\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"open","template":"","meta":{"footnotes":""},"class_list":["post-1163","page","type-page","status-publish","hentry","entry"],"_links":{"self":[{"href":"https:\/\/michaelnielsen.org\/blog\/wp-json\/wp\/v2\/pages\/1163","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/michaelnielsen.org\/blog\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/michaelnielsen.org\/blog\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/michaelnielsen.org\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/michaelnielsen.org\/blog\/wp-json\/wp\/v2\/comments?post=1163"}],"version-history":[{"count":9,"href":"https:\/\/michaelnielsen.org\/blog\/wp-json\/wp\/v2\/pages\/1163\/revisions"}],"predecessor-version":[{"id":1174,"href":"https:\/\/michaelnielsen.org\/blog\/wp-json\/wp\/v2\/pages\/1163\/revisions\/1174"}],"wp:attachment":[{"href":"https:\/\/michaelnielsen.org\/blog\/wp-json\/wp\/v2\/media?parent=1163"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}