<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Leverge — Insights</title><description>Field notes on building AI agents, RAG systems and LLM applications that hold up in production.</description><link>https://leverge.ai</link><language>en-us</language><managingEditor>hello@leverge.ai</managingEditor><webMaster>hello@leverge.ai</webMaster><item><title>How to Build an LLM Evaluation Set That Is Worth Trusting</title><link>https://leverge.ai/insights/building-an-llm-evaluation-set</link><guid isPermaLink="true">https://leverge.ai/insights/building-an-llm-evaluation-set</guid><description>The practical method — where the cases come from, how to adjudicate them, and how to validate a model-as-judge before you trust its numbers.</description><pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate><category>evaluation</category><category>evaluation</category><category>LLMOps</category><category>testing</category><category>quality</category></item><item><title>Why AI Agents Fail in Production</title><link>https://leverge.ai/insights/why-ai-agents-fail-in-production</link><guid isPermaLink="true">https://leverge.ai/insights/why-ai-agents-fail-in-production</guid><description>Six failure modes that account for almost every agent project we have been called in to rescue, and what actually prevents each one.</description><pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate><category>ai-agents</category><category>AI agents</category><category>production</category><category>evaluation</category><category>reliability</category></item></channel></rss>