llm-evaluation

Name: llm-evaluation - Agent Skill
Rating: 5.0 (26449 reviews)
Author: wshobson

26,449

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

Download Folder View on GitHub

CLI Install

Recommended

Use the skills CLI to install this skill with one command. Auto-detects all installed AI assistants.

Method 1 - skills CLI

npx skills i wshobson/agents/plugins/llm-application-dev/skills/llm-evaluation

Method 2 - openskills (supports sync & update)

npx openskills install wshobson/agents

Auto-detects Claude Code, Cursor, Codex CLI, Gemini CLI, and more. One install, works everywhere.

Installation Path

Download and extract to one of the following locations:

~/.claude/skills/llm-evaluation/

SKILL.md

Skill Instructions

Back

Run with Cloud Agent

No setup needed. Let our cloud agents run this skill for you.

Select Provider

Select Model

Claude Sonnet 4.5

$0.20/task

Best for coding tasks

No setup required

LLM Evaluation

Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

When to Use This Skill

Measuring LLM application performance systematically
Comparing different models or prompts
Detecting performance regressions before deployment
Validating improvements from prompt changes
Building confidence in production systems
Establishing baselines and tracking progress over time
Debugging unexpected model behavior

Core Evaluation Types

1. Automated Metrics

Fast, repeatable, scalable evaluation using computed scores.

Text Generation:

BLEU: N-gram overlap (translation)
ROUGE: Recall-oriented (summarization)
METEOR: Semantic similarity
BERTScore: Embedding-based similarity
Perplexity: Language model confidence

Classification:

Accuracy: Percentage correct
Precision/Recall/F1: Class-specific performance
Confusion Matrix: Error patterns
AUC-ROC: Ranking quality

Retrieval (RAG):

MRR: Mean Reciprocal Rank
NDCG: Normalized Discounted Cumulative Gain
Precision@K: Relevant in top K
Recall@K: Coverage in top K

2. Human Evaluation

Manual assessment for quality aspects difficult to automate.

Dimensions:

Accuracy: Factual correctness
Coherence: Logical flow
Relevance: Answers the question
Fluency: Natural language quality
Safety: No harmful content
Helpfulness: Useful to the user

3. LLM-as-Judge

Use stronger LLMs to evaluate weaker model outputs.

Approaches:

Pointwise: Score individual responses
Pairwise: Compare two responses
Reference-based: Compare to gold standard
Reference-free: Judge without ground truth

Quick Start

from dataclasses import dataclass
from typing import Callable
import numpy as np
 
@dataclass
class Metric

Automated Metrics Implementation

BLEU Score

from nltk.translate.bleu_score import sentence_bleu, SmoothingFunction
 
def calculate_bleu(reference: str, hypothesis: str, **kwargs) -> float:
    """Calculate BLEU score between reference and hypothesis."""
    smoothie = SmoothingFunction().method4
 
    return sentence_bleu(
        [reference.split()],
        hypothesis.split(),
        smoothing_function=smoothie
    )

ROUGE Score

from rouge_score import rouge_scorer
 
def calculate_rouge(reference: str, hypothesis: str, **kwargs) -> dict:
    """Calculate ROUGE scores."""
    scorer = rouge_scorer.RougeScorer(
        ['rouge1', 'rouge2', 'rougeL'],
        use_stemmer=True
    )

BERTScore

from bert_score import score
 
def calculate_bertscore(
    references: list[str],
    hypotheses: list[str],
    **kwargs
) -> dict:
    """Calculate BERTScore using pre-trained model."""
    P, R, F1 = score(
        hypotheses,
        references,
        lang='en',

Custom Metrics

def calculate_groundedness(response: str, context: str, **kwargs) -> float:
    """Check if response is grounded in provided context."""
    from transformers import pipeline
 
    nli = pipeline(
        "text-classification"

LLM-as-Judge Patterns

Single Output Evaluation

from anthropic import Anthropic
from pydantic import BaseModel, Field
import json
 
class QualityRating(BaseModel):
    accuracy: int = Field(ge

Pairwise Comparison

from pydantic import BaseModel, Field
from typing import Literal
 
class ComparisonResult(BaseModel):
    winner: Literal["A", "B", "tie"]
    reasoning: str
    confidence: int =

Reference-Based Evaluation

class ReferenceEvaluation(BaseModel):
    semantic_similarity: float = Field(ge=0, le=1)
    factual_accuracy: float = Field(ge=0

Human Evaluation Frameworks

Annotation Guidelines

from dataclasses import dataclass, field
from typing import Optional
 
@dataclass
class AnnotationTask:
    """Structure for human annotation task."""
    response: str
    question: str
    context: Optional[str]

Inter-Rater Agreement

from sklearn.metrics import cohen_kappa_score
 
def calculate_agreement(
    rater1_scores: list[int],
    rater2_scores: list[int]
) -> dict:
    """Calculate inter-rater agreement."""
    kappa = cohen_kappa_score(rater1_scores, rater2_scores)
 
    if kappa < 0:
        interpretation

A/B Testing

Statistical Testing Framework

from scipy import stats
import numpy as np
from dataclasses import dataclass, field

Regression Testing

Regression Detection

from dataclasses import dataclass
 
@dataclass
class RegressionResult:
    metric: str
    baseline: float
    current: float
    change: float
    is_regression:

LangSmith Evaluation Integration

from langsmith import Client
from langsmith.evaluation import evaluate, LangChainStringEvaluator
 
# Initialize LangSmith client
client = Client()
 
# Create dataset
dataset = client.create_dataset("qa_test_cases")
client.create_examples(
    inputs

Benchmarking

Running Benchmarks

from dataclasses import dataclass
import numpy as np
 
@dataclass
class BenchmarkResult:
    metric: str
    mean: float
    std: float

Resources

Best Practices

Multiple Metrics: Use diverse metrics for comprehensive view
Representative Data: Test on real-world, diverse examples
Baselines: Always compare against baseline performance
Statistical Rigor: Use proper statistical tests for comparisons
Continuous Evaluation: Integrate into CI/CD pipeline
Human Validation: Combine automated metrics with human judgment
Error Analysis: Investigate failures to understand weaknesses
Version Control: Track evaluation results over time

Common Pitfalls

Single Metric Obsession: Optimizing for one metric at the expense of others
Small Sample Size: Drawing conclusions from too few examples
Data Contamination: Testing on training data
Ignoring Variance: Not accounting for statistical uncertainty
Metric Mismatch: Using metrics not aligned with business goals
Position Bias: In pairwise evals, randomize order
Overfitting Prompts: Optimizing for test set instead of real use

return Metric("accuracy", calculate_accuracy)

return Metric("bleu", calculate_bleu)

return Metric("bertscore", calculate_bertscore)

def custom(name: str, fn: Callable):

return Metric(name, fn)

class EvaluationSuite:

def __init__(self, metrics: list[Metric]):

self.metrics = metrics

async def evaluate(self, model, test_cases: list[dict]) -> dict:

results = {m.name: [] for m in self.metrics}

for test in test_cases:

prediction = await model.predict(test["input"])

for metric in self.metrics:

prediction=prediction,

reference=test.get("expected"),

context=test.get("context")

results[metric.name].append(score)

"metrics": {k: np.mean(v) for k, v in results.items()},

"raw_scores": results

suite = EvaluationSuite([

Metric.custom("groundedness", check_groundedness)

"input": "What is the capital of France?",

"expected": "Paris",

"context": "France is a country in Europe. Paris is its capital."

results = await suite.evaluate(model=your_model, test_cases=test_cases)

scorer.score(reference, hypothesis)

'rouge1': scores['rouge1'].fmeasure,

'rouge2': scores['rouge2'].fmeasure,

'rougeL': scores['rougeL'].fmeasure

'microsoft/deberta-xlarge-mnli'

'precision': P.mean().item(),

'recall': R.mean().item(),

'f1': F1.mean().item()

model="microsoft/deberta-large-mnli"

result = nli(f"{context} [SEP] {response}")[0]

# Return confidence that response is entailed by context

return result['score'] if result['label'] == 'ENTAILMENT' else 0.0

def calculate_toxicity(text: str, **kwargs) -> float:

"""Measure toxicity in generated text."""

from detoxify import Detoxify

results = Detoxify('original').predict(text)

return max(results.values()) # Return highest toxicity score

def calculate_factuality(claim: str, sources: list[str], **kwargs) -> float:

"""Verify factual claims against sources."""

from transformers import pipeline

nli = pipeline("text-classification", model="facebook/bart-large-mnli")

for source in sources:

result = nli(f"{source}</s></s>{claim}")[0]

if result['label'] == 'entailment':

scores.append(result['score'])

return max(scores) if scores else 0.0

"Factual correctness"

helpfulness: int = Field(ge=1, le=10, description="Answers the question")

clarity: int = Field(ge=1, le=10, description="Well-written and understandable")

reasoning: str = Field(description="Brief explanation")

async def llm_judge_quality(

"""Use Claude to judge response quality."""

client = Anthropic()

system = """You are an expert evaluator of AI responses.

Rate responses on accuracy, helpfulness, and clarity (1-10 scale).

Provide brief reasoning for your ratings."""

prompt = f"""Rate the following response:

Question: {question}

{f'Context: {context}' if context else ''}

Response: {response}

Provide ratings in JSON format:

"helpfulness": <1-10>,

"reasoning": "<brief explanation>"

message = client.messages.create(

model="claude-sonnet-4-5",

messages=[{"role": "user", "content": prompt}]

return QualityRating(**json.loads(message.content[0].text))

async def compare_responses(

) -> ComparisonResult:

"""Compare two responses using LLM judge."""

client = Anthropic()

prompt = f"""Compare these two responses and determine which is better.

Question: {question}

Response A: {response_a}

Response B: {response_b}

Consider accuracy, helpfulness, and clarity.

"winner": "A" or "B" or "tie",

"reasoning": "<explanation>",

"confidence": <1-10>

message = client.messages.create(

model="claude-sonnet-4-5",

messages=[{"role": "user", "content": prompt}]

return ComparisonResult(**json.loads(message.content[0].text))

completeness: float = Field(ge=0, le=1)

async def evaluate_against_reference(

) -> ReferenceEvaluation:

"""Evaluate response against gold standard reference."""

client = Anthropic()

prompt = f"""Compare the response to the reference answer.

Question: {question}

Reference Answer: {reference}

Response to Evaluate: {response}

1. Semantic similarity (0-1): How similar is the meaning?

2. Factual accuracy (0-1): Are all facts correct?

3. Completeness (0-1): Does it cover all key points?

4. List any specific issues or errors.

"semantic_similarity": <0-1>,

"factual_accuracy": <0-1>,

"completeness": <0-1>,

"issues": ["issue1", "issue2"]

message = client.messages.create(

model="claude-sonnet-4-5",

messages=[{"role": "user", "content": prompt}]

return ReferenceEvaluation(**json.loads(message.content[0].text))

def get_annotation_form(self) -> dict:

"question": self.question,

"context": self.context,

"response": self.response,

"description": "Is the response factually correct?"

"description": "Does it answer the question?"

"description": "Is it logically consistent?"

"factual_error": False,

"hallucination": False,

"unsafe_content": False

interpretation = "Slight"

interpretation = "Fair"

interpretation = "Moderate"

interpretation = "Substantial"

interpretation = "Almost Perfect"

"interpretation": interpretation

variant_a_name: str = "A"

variant_b_name: str = "B"

variant_a_scores: list[float] = field(default_factory=list)

variant_b_scores: list[float] = field(default_factory=list)

def add_result(self, variant: str, score: float):

"""Add evaluation result for a variant."""

self.variant_a_scores.append(score)

self.variant_b_scores.append(score)

def analyze(self, alpha: float = 0.05) -> dict:

"""Perform statistical analysis."""

a_scores = np.array(self.variant_a_scores)

b_scores = np.array(self.variant_b_scores)

t_stat, p_value = stats.ttest_ind(a_scores, b_scores)

# Effect size (Cohen's d)

pooled_std = np.sqrt((np.std(a_scores)**2 + np.std(b_scores)**2) / 2)

cohens_d = (np.mean(b_scores) - np.mean(a_scores)) / pooled_std

"variant_a_mean": np.mean(a_scores),

"variant_b_mean": np.mean(b_scores),

"difference": np.mean(b_scores) - np.mean(a_scores),

"relative_improvement": (np.mean(b_scores) - np.mean(a_scores)) / np.mean(a_scores),

"statistically_significant": p_value < alpha,

"cohens_d": cohens_d,

"effect_size": self._interpret_cohens_d(cohens_d),

"winner": self.variant_b_name if np.mean(b_scores) > np.mean(a_scores) else self.variant_a_name

def _interpret_cohens_d(d: float) -> str:

"""Interpret Cohen's d effect size."""

class RegressionDetector:

def __init__(self, baseline_results: dict, threshold: float = 0.05):

self.baseline = baseline_results

self.threshold = threshold

def check_for_regression(self, new_results: dict) -> dict:

"""Detect if new results show regression."""

for metric in self.baseline.keys():

baseline_score = self.baseline[metric]

new_score = new_results.get(metric)

if new_score is None:

# Calculate relative change

relative_change = (new_score - baseline_score) / baseline_score

# Flag if significant decrease

is_regression = relative_change < -self.threshold

regressions.append(RegressionResult(

baseline=baseline_score,

change=relative_change,

"has_regression": len(regressions) > 0,

"regressions": regressions,

"summary": f"{len(regressions)} metric(s) regressed"

outputs=[{"answer": a} for a in expected_answers],

dataset_id=dataset.id

LangChainStringEvaluator("qa"), # QA correctness

LangChainStringEvaluator("context_qa"), # Context-grounded QA

LangChainStringEvaluator("cot_qa"), # Chain-of-thought QA

async def target_function(inputs: dict) -> dict:

result = await your_chain.ainvoke(inputs)

return {"answer": result}

experiment_results = await evaluate(

evaluators=evaluators,

experiment_prefix="v1.0.0",

metadata={"model": "claude-sonnet-4-5", "version": "1.0.0"}

print(f"Mean score: {experiment_results.aggregate_metrics['qa']['mean']}")

class BenchmarkRunner:

def __init__(self, benchmark_dataset: list[dict]):

self.dataset = benchmark_dataset

async def run_benchmark(

metrics: list[Metric]

) -> dict[str, BenchmarkResult]:

"""Run model on benchmark and calculate metrics."""

results = {metric.name: [] for metric in metrics}

for example in self.dataset:

# Generate prediction

prediction = await model.predict(example["input"])

# Calculate each metric

for metric in metrics: