开源 LLM 工程平台:追踪、评测、提示词管理与指标。
🧪 评测·观测·安全
模型评测、提示词测试、可观测性与护栏红队
测试提示词、模型与 RAG 的评测与红队工具,CLI 与 CI 友好。
大模型少样本评测框架,Open LLM Leaderboard 背后的工具。
Open-source AI penetration testing tool to find and fix your app’s vulnerabilities.
用众多可复用的提示词模式增强人类能力的开源框架。
An AI prompt optimizer for writing better prompts and getting better AI results.
817 structured cybersecurity skills for AI agents · Mapped to 6 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND, NIST AI RMF & MITRE F3 (Fight Fraud) · agentskills.io standard · Works with Claude Code, GitHub
调试、评测与监控 LLM 应用的开源平台。
Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, security risks, prompt injection, data exfiltration, and supply-chain risks in Claude Code, Codex, and MCP skills before you install them.
评测大模型与系统的框架及基准注册表。
LLM 评测框架,像 Pytest 一样为模型写测试。
评测与优化 RAG 流水线的框架。
LangGPT: Empowering everyone to become a prompt expert! 🚀 📌 结构化提示词(Structured Prompt)提出者 📌 元提示词(Meta-Prompt)发起者 📌 最流行的提示词落地范式 | Language of GPT The pioneering framework for structured & meta-prompt design 10,000+ ⭐ |
AI 可观测性与评测平台,追踪与调试 LLM 应用。
An open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data. Supports NLP, pattern matching, and customizable pipelines.
Build high-quality LLM apps - from prototyping, testing to production deployment and monitoring.
LLM 漏洞扫描器,探测幻觉、泄露与越狱风险。
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
上海 AI 实验室的大模型评测平台。
为 LLM 添加护栏,校验输入输出并生成结构化数据。
基于 OpenTelemetry 的 LLM 应用观测。
为 LLM 对话系统添加可编程护栏。
Fit interpretable models. Explain blackbox machine learning.
Optimize prompts, code, and more with AI-powered Reflective Optimization