将杂乱的文档集合转换为结构化行数据,使用 DocETL
Define repeatable extraction pipelines that pull fields from large document collections, normalize outputs, and audit failures across the corpus.
Prerequisites
Python 3.10+, DocETL, document corpus, extraction configuration
Installation
Use the upstream install or setup path that matches your environment:
- Use Docker (recommended for quick start): make docker
- pip install docetl
- Run Docker:
- make docker
Requirements and caveats from upstream:
- A Python package for running production pipelines from the command line or Python code
-
2. 📦 Python Package (For Production Use)
- If you want to use DocETL as a Python package:
Basic usage or getting-started notes:
-
🚀 Getting Started
-
DocWrangler is hosted at docetl.org/playground. But to run the playground locally, you can either:
-
OpenAI API key
-
Source: https://github.com/1991513ccie-png/skills
-
Extracted from upstream docs: https://raw.githubusercontent.com/ucbepic/docetl/HEAD/README.md
Documentation
- https://docetl.org/
微信扫一扫