Skip to content

添加 robots.txt 与许可证 - #1

Open
tangyuan0821 wants to merge 3 commits into
win12-online:mainfrom
tangyuan0821:main
Open

添加 robots.txt 与许可证#1
tangyuan0821 wants to merge 3 commits into
win12-online:mainfrom
tangyuan0821:main

Conversation

@tangyuan0821

Copy link
Copy Markdown
Member

如题

@tjy-gitnub tjy-gitnub left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

看不懂awa

Comment thread robots.txt
Comment on lines +216 to +220
User-agent: archive.org_bot
Disallow: /

User-agent: ia_archiver-web.archive.org
Disallow: /

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

为什么要禁止 Internet Archive :<

Comment thread robots.txt
Comment on lines +8 to +35
#Following spiders are allowed:
#Applebot
#Bingbot
#AdIdxBot
#BingPreview
#Googlebot
#Googlebot-mobile
#Googlebot-Image
#Googlebot-Video
#DuckDuckBot
#Sogou web spider
#Sogou inst spider
#Sogou spider
#Sogou wap spider
#360Spider
#360Spider-Image
#360Spider-Video
#Baiduspider
#Baiduspider-image
#Baiduspider-video
#Baiduspider-news
#Yisouspider
#ByteSpider
#Yandex
#PetalBot
#Slurp
#Yeti
#Other spiders are disallowed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

你真想允许某些bot的话直接按照 https://www.robotstxt.org/robotstxt.html 里的“To allow a single robot”示例写就可以

# 你允许的bot
User-agent: Google
Disallow:

# 禁止其余的
User-agent: *
Disallow: /

Comment thread robots.txt
Comment on lines +120 to +451
User-agent: MJ12bot
Disallow: /

User-agent: Mediapartners-Google*
Disallow: /

User-agent: UbiCrawler
Disallow: /

User-agent: DOC
Disallow: /

User-agent: Zao
Disallow: /

User-agent: sitecheck.internetseer.com
Disallow: /

User-agent: Zealbot
Disallow: /

User-agent: MSIECrawler
Disallow: /

User-agent: SiteSnagger
Disallow: /

User-agent: WebStripper
Disallow: /

User-agent: WebCopier
Disallow: /

User-agent: Fetch
Disallow: /

User-agent: Offline Explorer
Disallow: /

User-agent: Teleport
Disallow: /

User-agent: TeleportPro
Disallow: /

User-agent: WebZIP
Disallow: /

User-agent: linko
Disallow: /

User-agent: HTTrack
Disallow: /

User-agent: Microsoft.URL.Control
Disallow: /

User-agent: Xenu
Disallow: /

User-agent: larbin
Disallow: /

User-agent: libwww
Disallow: /

User-agent: ZyBORG
Disallow: /

User-agent: Download Ninja
Disallow: /

User-agent: fast
Disallow: /

User-agent: wget
Disallow: /

User-agent: grub-client
Disallow: /

User-agent: k2spider
Disallow: /

User-agent: NPBot
Disallow: /

User-agent: WebReaper
Disallow: /

User-agent: Browsershots
Disallow: /

User-agent: ia_archiver
Disallow: /

User-agent: archive.org_bot
Disallow: /

User-agent: ia_archiver-web.archive.org
Disallow: /

User-agent: AmazonBot
Disallow: /

User-agent: anthropic-ai
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: ChatGLM
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-Web
Disallow: /

User-agent: cohere-ai
Disallow: /

User-agent: Diffbot
Disallow: /

User-agent: FacebookBot
Disallow: /

User-agent: FriendlyCrawler
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: GoogleOther
Disallow: /

User-agent: GoogleOther-Image
Disallow: /

User-agent: GoogleOther-Video
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: ImagesiftBot
Disallow: /

User-agent: img2dataset
Disallow: /

User-agent: omgili
Disallow: /

User-agent: omgilibot
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: SkyworkSpider
Disallow: /

User-agent: Timpibot
Disallow: /

User-agent: YouBot
Disallow: /

User-agent: AI2Bot
Disallow: /

User-agent: Ai2Bot-Dolma
Disallow: /

User-agent: aiHitBot
Disallow: /

User-agent: Andibot
Disallow: /

User-agent: bedrockbot
Disallow: /

User-agent: Brightbot 1.0
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: cohere-training-data-crawler
Disallow: /

User-agent: Cotoyogi
Disallow: /

User-agent: Crawlspace
Disallow: /

User-agent: DuckAssistBot
Disallow: /

User-agent: EchoboxBot
Disallow: /

User-agent: facebookexternalhit
Disallow: /

User-agent: Factset_spyderbot
Disallow: /

User-agent: FirecrawlAgent
Disallow: /

User-agent: Google-CloudVertexBot
Disallow: /

User-agent: iaskspider/2.0
Disallow: /

User-agent: ICC-Crawler
Disallow: /

User-agent: ISSCyberRiskCrawler
Disallow: /

User-agent: Kangaroo Bot
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: meta-externalfetcher
Disallow: /

User-agent: MistralAI-User/1.0
Disallow: /

User-agent: MyCentralAIScraperBot
Disallow: /

User-agent: NovaAct
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: Operator
Disallow: /

User-agent: PanguBot
Disallow: /

User-agent: Panscient
Disallow: /

User-agent: panscient.com
Disallow: /

User-agent: Perplexity-User
Disallow: /

User-agent: PhindBot
Disallow: /

User-agent: Poseidon Research Crawler
Disallow: /

User-agent: QualifiedBot
Disallow: /

User-agent: QuillBot
Disallow: /

User-agent: quillbot.com
Disallow: /

User-agent: SBIntuitionsBot
Disallow: /

User-agent: Scrapy
Disallow: /

User-agent: SemrushBot
Disallow: /

User-agent: SemrushBot-BA
Disallow: /

User-agent: SemrushBot-CT
Disallow: /

User-agent: SemrushBot-OCOB
Disallow: /

User-agent: SemrushBot-SI
Disallow: /

User-agent: SemrushBot-SWA
Disallow: /

User-agent: Sidetrade indexer bot
Disallow: /

User-agent: TikTokSpider
Disallow: /

User-agent: VelenPublicWebCrawler
Disallow: /

User-agent: Webzio-Extended
Disallow: /

User-agent: wpbot
Disallow: /

User-agent: YandexAdditional
Disallow: /

User-agent: YandexAdditionalBot
Disallow: /

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

其实...没必要列举这么多的,如果你想禁止AI爬取的话可以只写常见的:

(摘自Cloudflare的「阻止训练」配置下的robots.txt)

User-agent: Amazonbot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CloudflareBrowserRenderingCrawler
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: meta-externalagent
Disallow: /

如果你想使用 Content Signals 的话,可在最上方添加这些内容:

(也是取自Cloudflare)

# As a condition of accessing this website, you agree to abide by the following
# content signals:

# (a)  If a Content-Signal = yes, you may collect content for the corresponding
#      use.
# (b)  If a Content-Signal = no, you may not collect content for the
#      corresponding use.
# (c)  If the website operator does not include a Content-Signal for a
#      corresponding use, the website operator neither grants nor restricts
#      permission via Content-Signal with respect to the corresponding use.

# The content signals and their meanings are:

# search:   building a search index and providing search results (e.g., returning
#           hyperlinks and short excerpts from your website's contents). Search does not
#           include providing AI-generated search summaries.
# ai-input: inputting content into one or more AI models (e.g., retrieval
#           augmented generation, grounding, or other real-time taking of content for
#           generative AI search answers).
# ai-train: training or fine-tuning AI models.
# use:      how AI systems may consume the content (immediate, reference, or full).

# ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF
# RIGHTS UNDER ARTICLE 4 OF THE EUROPEAN UNION DIRECTIVE 2019/790 ON COPYRIGHT
# AND RELATED RIGHTS IN THE DIGITAL SINGLE MARKET.

User-agent: *
Content-Signal: search=yes,ai-train=no,use=reference
Allow: /

不过,我们真的需要防止AI爬取吗...

@tangyuan0821

Copy link
Copy Markdown
Member Author

@lingbopro 感觉都是一些无伤大雅的小问题诶()

@lingbopro

Copy link
Copy Markdown
Member

@lingbopro 感觉都是一些无伤大雅的小问题诶()

小在哪

为什么要禁用 Internet Archive

为什么要全列出来

只禁用一些常见的就够了

更何况我们或许压根就不需要反爬...

@Iamliuxiaozhen Iamliuxiaozhen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这么写rebots.txt够了吧,我们貌似不需要阻止任何爬虫

User-agent: *
Allow: /

@lingbopro

Copy link
Copy Markdown
Member

这么写rebots.txt够了吧,我们貌似不需要阻止任何爬虫

User-agent: *
Allow: /

这样的话那还有写robots.txt的必要吗……(

@Iamliuxiaozhen

Copy link
Copy Markdown
Member

这样的话那还有写robots.txt的必要吗……(

实际上是有的,让他们大胆爬()

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants