Seenthis
•
 
Identifiants personnels
  • [mot de passe oublié ?]

 
  • #c
  • #cr
  • #cra
  • #crawl
RSS: #crawler

#crawler

  • #crawlers
  • @biggrizzly
    BigGrizzly @biggrizzly CC BY-NC-SA 10/09/2026
    2
    @raphael4
    @rastapopoulos
    2

    Xamanismo Coletivo - Mastodon
    ▻https://hachyderm.io/@eliasulrich/117234902637900790

    "One guy in #Sweden built a #searchengine to fight Google, and it works.

    It’s called #MarginaliaSearch.

    It runs its own #crawler and builds its own #index instead of borrowing Bing’s. It has no ads, no investors, and no loans.

    What it does differently: it ranks for text-heavy, non-commercial pages. Personal blogs. Old university pages.

    The weird corners SEO strangled. Every result tells you whether the page uses affiliate links and JavaScript, and you can filter them out.

    There’s an “explore” mode that just shows you random sites from the index. It’s open source under AGPL, so you can host your own copy.

    It’s keyword-based, so don’t type a full question at it. Type two nouns and see where you land. Every #searchengine now shows you the same twelve #monetizedpages."

    ▻https://marginalia-search.com

    BigGrizzly @biggrizzly CC BY-NC-SA
    Écrire un commentaire
  • @biggrizzly
    BigGrizzly @biggrizzly CC BY-NC-SA 16/08/2025

    #Codeberg - Mastodon
    ▻https://social.anoxinon.de/@Codeberg/115033790447125787

    It seems like the #AI #crawlers learned how to solve the #Anubis #challenges. Anubis is a tool hosted on our #infrastructure that requires #browsers to do some heavy #computation before accessing Codeberg again. It really saved us tons of nerves over the past months, because it saved us from manually maintaining #blocklists to having a working detection for “real browsers” and “AI crawlers”.

    BigGrizzly @biggrizzly CC BY-NC-SA
    Écrire un commentaire
  • @oanth_rss
    oAnth_RSS @oanth_rss CC BY 1/07/2025
    2
    @gao_tumbuktu
    @02myseenthis01
    2

    Cloudflare declares war on AI crawlers - and the stakes couldn’t be higher

    via ▻https://diasp.eu/p/17724248

    ▻https://www.zdnet.com/article/cloudflare-declares-war-on-ai-crawlers-and-the-stakes-couldnt-be-higher

    #computers #security #technology #news #education #updates #tech #analysis #research

    oAnth_RSS @oanth_rss CC BY
    • @02myseenthis01
      oAnth @02myseenthis01 CC BY 2/07/2025

      Get out of my website! Cloudflare, one of the world’s largest Internet infrastructure providers, has begun blocking AI web crawlers by default.

      Written by Steven Vaughan-Nichols, Senior Contributing Editor

      July 1, 2025 at 1:11 p.m. PT

      The major Internet Content Delivery Network (CDN), Cloudflare, has declared war on AI companies. Starting July 1, Cloudflare now blocks by default AI web crawlers accessing content from your websites without permission or compensation.

      The change addresses a real problem. My own small site, where I track all my stories, Practical Technology, has been slowed dramatically at times by AI crawlers. It’s not just me. Numerous website owners have reported that AI crawlers, such as OpenAI’s GPTBot and Anthropic’s ClaudeBot, generate massive volumes of automated requests that clog up websites so they’re as slow as sludge. GoogleBot alone reports that the cloud-hosting service Vercel bombards the sites it hosts with over 4.5 billion requests a month.

      These AI bots often crawl sites far more aggressively than traditional search engine crawlers. They sometimes revisit the same pages every few hours or even hit sites with hundreds of requests per second. While the AI companies deny that their bots are to blame, the evidence tells a different story.

      Thus, on behalf of its two million-plus customers, 20% of the web, Cloudflare now blocks #AI_crawlers. For any new website signing up for its services, AI crawlers will be automatically blocked from accessing its content unless the site owner grants explicit permission. Additionally, Cloudflare promises to detect “shadow” #scrapers — bots that attempt to evade detection — by using behavioral analysis and machine learning. What’s good for the AI goose is good for the gander.

      This move reverses the previous status quo, where website owners had to opt out of AI crawling. Now, blocking is the default, and AI vendors must request access and clarify their intentions, whether for model training, search, or other uses, before they’re allowed in.

      This change arises not only because of frustrated website owners. Numerous publishing companies, such as The Associated Press, Condé Nast, and ZDNET’s own parent company, Ziff Davis, are frustrated that #AI companies have been “strip mining” the web for content. All too often, this has been done without compensation or consent, and sometimes, ignoring standard protocols like robots.txt that are meant to block #crawlers.

      (Disclosure: Ziff Davis, ZDNET’s parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)

      Moreover, recent court cases have ruled in favor of Meta and Anthropic, finding that their use of copyrighted works was legal under the doctrine of fair use. Needless to say, writers, artists, and publishers don’t like this one bit. Publishers are still worried that the federal government will give AI free rein to do as it wants with their content. AI powerhouses such as #OpenAI and Google are continuing to lobby the government to classify AI training on copyrighted data as fair use.

      It’s also worth noting that after the Copyright Office released a pre-publication version of its 108-page #copyright and #AI report, which struck a middle ground by supporting both of these world-class industries that contribute so much to our economic and cultural advancement. However, it added that while some generative AI probably constitutes a “transformative” use, the mass scraping of all data did not qualify as fair use. The next day, the Trump administration fired the head of the Copyright Office and replaced her with an attorney with no prior experience in #copyright_law.

      Given all this, it’s no wonder that publishers sought an ally in technology.

      As Cloudflare CEO Matthew Prince said in a statement, its new policy is meant to “give publishers the control they deserve and build a new economic model that works for everyone—creators, consumers, tomorrow’s AI founders, and the future of the web itself.”

      To complement the move to block AI crawlers, Cloudflare has also launched its “Pay Per Crawl” program. This enables publishers to set their own rates for AI companies that want to scrape their content.

      This system is currently in private beta and aims to create a framework where AI firms can pay for access, or be denied if they refuse. Technically, this will be done by dusting off an old, mostly unused web server response, HTTP 402, which responds with a “Payment Required” error message. This means it should be simple to implement and compatible with existing websites and their infrastructure.

      Overall, this is a big deal. Thanks to Cloudflare powering such a large portion of the internet, a significant amount of web content could become inaccessible to AI companies unless they negotiate access or pay licensing fees. As Nicholas Thompson, CEO of The Atlantic, noted, “Until now, AI companies have not needed to pay for content licenses because they could simply take it without repercussions. Now they will need to negotiate.”

      To this point, most AI companies have been actively against paying for content. As Sir Nick Clegg, former deputy UK Prime Minister and Meta executive, said recently, merely asking artists’ permission before they scrape copyrighted content will “basically kill the AI industry.”

      Cloudflare’s new policy is a direct response to this approach and the increasing volume and intrusiveness of AI crawlers that have come with it. It’s also an attempt to stop the siphoning of traffic that would otherwise go to publishers.

      Since the rise of AI, traffic to news sites has plunged. For example, Business Insider’s traffic dropped by over half, 55% from April 2022 to April 2025. Left unchecked, Thompson recently predicted that, thanks to AI, the Atlantic staff should expect traffic from Google to drop to zero.

      [...]

      oAnth @02myseenthis01 CC BY
    • @02myseenthis01
      oAnth @02myseenthis01 CC BY 2/07/2025

      #AI : aspects évidents du droit d’ #auteur, de la protection contre la #copie ainsi que des intentions et fréquences d’ #accès et de la #monétisation

      oAnth @02myseenthis01 CC BY
    Écrire un commentaire
  • @b_b
    b_b @b_b PUBLIC DOMAIN 22/05/2025
    3
    @biggrizzly
    @arno
    @olaf
    3

    Improved ways to operate a rude #crawler marginalia.nu
    ▻https://www.marginalia.nu/log/a_115_rude_crawler

    Tech news is abuzz with rude AI crawlers that forge their user-agent and ignore robots.txt. In my opinion, if this is all the AI startups can muster, they’re losing their touch. wget can do this. You need to up your game, get that crawler really rolling coal. Flagrant disregard for externalities is an important signal to the investors that your AI startup is the one.

    In that spirit, here are some advanced tips on how to be a much worse netizen.

    #satire #bots #ia #botnet

    En lien avec ▻https://seenthis.net/messages/1104052

    b_b @b_b PUBLIC DOMAIN
    Écrire un commentaire
  • @aurelieng
    aurelieng @aurelieng via RSS CC BY 9/03/2025

    Mise à mal des forges Git par les indexeurs d’IA - Infrastructure - Forum du collectif CHATONS
    ►https://forum.chatons.org/t/mise-a-mal-des-forges-git-par-les-indexeurs-dia/7086

    — Permalink

    #LLMs #generativeai #copilot #training #web #crawlers

    aurelieng @aurelieng via RSS CC BY
    Écrire un commentaire
  • @cy_altern
    cy_altern @cy_altern CC BY-SA 23/04/2022
    1
    @hellodoc
    1

    « Disparition » de sites pirates : que s’est-il passé avec DuckDuckGo ? - Numerama
    ▻https://www.numerama.com/tech/926821-disparition-de-sites-pirates-que-sest-il-passe-avec-duckduckgo.html

    #DuckDuckGo #crawler #indexation #censure #filtrage

    cy_altern @cy_altern CC BY-SA
    Écrire un commentaire
  • @cy_altern
    cy_altern @cy_altern CC BY-SA 23/07/2020

    internetarchive/heritrix3: Heritrix is the Internet Archive’s open-source, extensible, web-scale, archival-quality web crawler project.
    ▻https://github.com/internetarchive/heritrix3

    Heritrix is the Internet Archive’s open-source, extensible, web-scale, archival-quality web crawler project.

    – Le wiki de documentation: ▻https://github.com/internetarchive/heritrix3/wiki
    – téléchargement: ▻http://builds.archive.org/maven2/org/archive/heritrix/heritrix

    #heritrix #crawler #aspirateur_site #internetarchive

    cy_altern @cy_altern CC BY-SA
    Écrire un commentaire
  • @mr_cerbere
    Mr Cerbere @mr_cerbere 12/03/2019

    Olivier PAPON ? sur Twitter : «  ?NEW ? ▻https://t.co/jO0w0gPfrT devient aussi un crawler ?️ surpuissant : 1000 urls/sec et 5M de pages/site ?. Une belle liste de features à découvrir : historique des crawls, codes http, profondeur, urls crawlées/bloquées... #SEO #crawler. Enjoy and please RT ? ?… ▻https://t.co/arKXrU1pCG »
    ▻https://twitter.com/seolyzer_io/status/1105034807069806592

    Mr Cerbere @mr_cerbere
    Écrire un commentaire
  • @bloginfo
    bloginfo @bloginfo CC BY-NC-ND 7/02/2016

    Les #Bots, #Spiders, #Crawlers à autoriser sur votre site Web
    ▻http://www.dsfc.net/internet/moteurs-internet/bots-spiders-crawlers-a-autoriser-sur-votre-site-web

    http://www3.pictures.zimbio.com/gi/Missy+Franklin+2012+T+Winter+National+Championships+9kIpFgITJlxl.jpg

    Contrairement à certains SEO, j’ai toujours considéré qu’il y avait un intérêt à être indexé dans des #Moteurs de recherche « mineurs ».

    #.htaccess #Formateur_Apache #Formateur_Référencement_naturel #Formateur_SEO #Moteurs_de_recherche #Search_Engines

    bloginfo @bloginfo CC BY-NC-ND
    Écrire un commentaire
  • @cy_altern
    cy_altern @cy_altern CC BY-SA 1/11/2015
    1
    @fil
    1

    Heritrix - Heritrix - IA Webteam Confluence
    ▻https://webarchive.jira.com/wiki/display/Heritrix/Heritrix

    Heritrix is the Internet Archive’s open-source, extensible, web-scale, archival-quality web #crawler project. Heritrix (sometimes spelled heretrix, or misspelled or mis-said as heratrix/heritix/ heretix/heratix) is an archaic word for heiress (woman who inherits). Since our crawler seeks to collect and preserve the digital artifacts of our culture for the benefit of future researchers and generations, this name seemed apt.

    #spider #wget #achive #aspirateur

    cy_altern @cy_altern CC BY-SA
    • @fil
      Fil @fil 1/11/2015

      #archivage_militant fais-moi signe si tu réussis à l’installer et à l’utiliser ?

      Fil @fil
    Écrire un commentaire
  • @cy_altern
    cy_altern @cy_altern CC BY-SA 4/01/2015
    1
    @spip
    1

    Detects a few common Search Bots
    ▻https://gist.github.com/ScottPhillips/2904459

    un script #php simple pour la #détection des robots d’indexation. A utiliser avec un preg_match("/$crawlers_names/i", $user_agent) 1 pour éviter le foreach

    #bot #crawler #robot

    cy_altern @cy_altern CC BY-SA
    • @ben
      Ben @ben CC BY-NC 4/01/2015

      j’avais repéré cette liste aussi ▻https://github.com/YOURLS/dont-log-bots/blob/master/plugin.php#L21 ( en faisant une recherche sur github sur l’un des bots)

      Ben @ben CC BY-NC
    • @b_b
      b_b @b_b PUBLIC DOMAIN 2/03/2015
      @cy_altern @ben @nicod_

      ... plop :)

      @cy_altern @ben @nicod_ ça serait pas intéressant de compléter notre liste dans l’écran de sécurité à partir des deux citées ici ?

      b_b @b_b PUBLIC DOMAIN
    • @cy_altern
      cy_altern @cy_altern CC BY-SA 5/03/2015

      après compilation et dédoublonnage des 2 listes proposées, ça donnerait le code suivant : ▻http://spip.pastebin.fr/39305
      N’est ce pas un peu trop gros comme expression régulière pour un preg_match() ?

      cy_altern @cy_altern CC BY-SA
    • @kent1
      kent1 @kent1 ART LIBRE 9/01/2018

      Au dessus on filtre déja bot|slurp|crawler|spider|webvac|yandex| donc si tu enlèves ceux qui matchent cela devrait aller mieux en théorie

      kent1 @kent1 ART LIBRE
    • @kent1
      kent1 @kent1 ART LIBRE 9/01/2018

      Version mise à jour : ▻http://spip.pastebin.fr/52828

      kent1 @kent1 ART LIBRE
    Écrire un commentaire
  • @sammyfisherjr
    SammyFisherJr @sammyfisherjr CC BY-NC-SA 18/11/2011

    Les pages perso de Free ne sont pas référencées sur Bing - Freenews : L’actualité des Freenautes - Toute l’actualité pour votre Freebox Revolution
    ►http://www.freenews.fr/spip.php?article11083

    Certains ont déjà pu le constater, depuis maintenant quelques mois, les pages personnelles hébergées sur Free.fr ne sont plus référencées sur le moteur de recherche Bing de Microsoft [...] le crawler (robot indexant les pages) de Bing a un comportement trop agressif, causant une surcharge bien trop importante sur les serveurs des pages perso

    #Bing #Free #crawler

    SammyFisherJr @sammyfisherjr CC BY-NC-SA
    Écrire un commentaire

Thèmes liés

  • #ai
  • #crawlers
  • #bing
  • #free
  • person: bing free
  • person: bing de microsoft