Seenthis
•
 
Identifiants personnels
  • [mot de passe oublié ?]

  • https://www.theverge.com

/24067997

  • ►/robots-txt-ai-text-file-web-crawlers-spiders
  • @hubertguillaud
    hubertguillaud @hubertguillaud CC BY 15/02/2024
    3
    @severo
    @monolecte
    @ericw
    3

    Excellente histoire du Robots.txt, ce fichier qui autorise ou non l’indexation. Excellente mise en perspective qui explique pourquoi beaucoup souhaitent interdire l’indexation pour l’IA, mais peine à identifier les robots dédiés au-delà de GPTbot : ►https://www.theverge.com/24067997/robots-txt-ai-text-file-web-crawlers-spiders

    Et le problème, c’est que les robots commencent à ne plus respecter les contraintes du fichier... #indexation

    hubertguillaud @hubertguillaud CC BY
    Écrire un commentaire
  • @supergeante
    Supergéante @supergeante 15/02/2024
    3
    @biggrizzly
    @simplicissimus
    @ericw
    3

    A propos de Robots.txt et des crawlers AI.

    The text file that runs the internet
    ▻https://duet-cdn.vox-cdn.com/thumbor/0x0:2040x1360/1200x800/filters:focal(1020x680:1021x681):format(webp)/cdn.vox-cdn.com/uploads/chorus_asset/file/25282612/246992_AI_at_Work_TXT_FILE_ECarter.png
    For decades, robots.txt governed the behavior of web crawlers. But as unscrupulous AI companies seek out more and more data, the basic social contract of the web is falling apart.
    ►https://www.theverge.com/24067997/robots-txt-ai-text-file-web-crawlers-spiders

    Supergéante @supergeante
    Écrire un commentaire