把整个网站转为适合 LLM 的 Markdown 或结构化数据的抓取 API。
📄 文档解析与抓取
OCR、PDF 转 Markdown、网页抓取:把数据喂给大模型
百度飞桨的 OCR 工具库,百余种语言、PP-OCR 与文档解析一体。
把 PDF 等文档高质量转换为 Markdown / JSON,面向大模型语料。
IBM 开源的文档解析工具,PDF / Office 转结构化数据。
微软出品,把 Office、PDF 等文件转为 Markdown 供 LLM 使用。
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/
为 LLM 设计的开源网页爬虫与抓取工具。
历史最悠久的开源 OCR 引擎。
快速准确地把 PDF 转为 Markdown 与 JSON。
Lightpanda: the headless browser designed for AI and automation
用大模型和图逻辑驱动的 Python 网页抓取库。
开箱即用的 80+ 语言 OCR 库。
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
DeepSeek 探索视觉压缩上下文的 OCR 模型。
AI2 的 PDF 线性化工具包,用视觉语言模型解析文档。
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
文档预处理库,为 LLM 提取并整理非结构化数据。
把任意 URL 转成 LLM 友好的输入,在前面加上 r.jina.ai 即可。
基于深度学习的文档文字识别库。
An automated document analyzer for Paperless-ngx using OpenAI API, Ollama, Deepseek-r1, Azure and all OpenAI API compatible Services to automatically analyze and tag your documents.
The go-to web for your AI coding agent — local-first search, fetch, crawl & research over MCP. No API keys, no cloud, $0/query. Public beta.
A system for agentic LLM-powered data processing and ETL
Open-source inference server and production cluster for all the models your agent needs.
모두 파싱해버리겠다 — HWP·HWPX·PDF·Office 문서를 Markdown으로. 양식 자동 채우기와 신구대조를 갖춘 CLI·MCP 서버 | Convert Korean documents (HWP, HWPX, PDF, Office) to Markdown — CLI and MCP server with form filling and diff