- OpenHands (83K★, open-source AI software engineer) — fixed automation cards silently dropping required integrations that can't be installed as MCP servers, with i18n and regression tests · Issue #16292 → PR #16324 (merged)
- DeepEval (17K★, LLM evaluation framework) — found and fixed nonexistent Claude model IDs that made every default-configured Anthropic evaluation fail with a 404 · Issue #2993 → PR #2994 (in review)
- DeepEval — found and fixed the judge's 1024-token output cap silently truncating responses on thinking-enabled Claude models, surfacing as cryptic invalid-JSON metric errors while thinking tokens were still billed · Issue #3042 → PR #3043 (in review)
- LiteLLM (56K★, LLM gateway) — traced a user's
/v1/v1/messages404 to the one Anthropic passthrough config missing trailing-/v1normalization, answered the report, and fixed it · Discussion #36651 → Issue #36956 → PR #36957 (in review)
Pinned Loading
-
agenteval
agenteval PublicReliability and audit-evidence testing for LLM agents - wrap any agent, assert behavior, measure determinism, check grounding, emit an audit-grade report.
TypeScript
-
-
litellm
litellm PublicForked from BerriAI/litellm
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthr…
Python
-
OpenHands
OpenHands PublicForked from OpenHands/OpenHands
🙌 OpenHands: AI-Driven Development
TypeScript
-
langchain-ai/langchain
langchain-ai/langchain PublicThe agent engineering platform.
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.



