多类型数据标注工具,支持文本、图像、音频与大模型评估。
🏷️ 数据与标注
数据集、标注、清洗与合成数据
Hugging Face 的数据集库,轻松访问与处理海量数据集。
A high-performance observability data pipeline.
GenBI (Generative BI) for AI agents, an open-source, governed text-to-SQL through an open context layer that turns natural-language questions into trusted dashboards, charts, and SQL across 20+ data sources, such as BigQ
业界常用的图像与视频标注工具。
The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.
The Context Platform for your Data and AI Stack
自动发现数据与标签中的问题。
可视化、管理并改进计算机视觉数据集。
Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.
开源文本标注工具。
🧙 Build, run, and manage data pipelines for integrating and transforming data.
Apache Beam is a unified programming model for Batch and Streaming data processing.
[SIGMOD'27] Easy Data Preparation with latest LLMs-based Operators and Pipelines.
面向 AI 的现代列式数据格式。
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
A system for quickly generating training data with weak supervision
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
面向 AI 工程师与领域专家的数据标注与协作平台。
TextAttack 🐙 is a Python framework for adversarial attacks, data augmentation, and model training in NLP https://textattack.readthedocs.io/en/master/
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.
Apache Ossie, industry wide specification effort to standardize how we exchange semantic metadata across analytics, AI and BI platforms, providing a vendor neutral, single source of truth for semantic data