0.6
- New Features
- Golden Datasets (New): Build evaluation datasets from real production traffic by promoting spans into a dataset.
Application.get_span_fields()discovers which attribute keys your spans carry, with coverage counts and any evaluator outputs recorded on them.Dataset.add_items_from_spans()resolves span references server-side, applies aFieldMapping, and writes one dataset item per span.Dataset.get_schema()introspects an existing dataset’s fields so a new mapping stays consistent with it. Writes are all-or-nothing: if any span cannot be resolved, nothing is written and the call returns422with aSpanNotFoundentry per unresolved span. Adds theSpanReference,FieldMapping,SearchFilter, andSearchScopemodels; structured span filtering reuses the existingQueryCondition,QueryRule, andOperatorTypemodels. - Experiment Trace Capture (New):
evaluate()now captures the OpenTelemetry spans your task emits and links them to the experiment item that produced them, with no configuration. Evaluators opt in to reading them by declaring asessionparameter, which receives the item’s captured spans and enables scoring on execution — tool call counts, latency, retries, span exception events — rather than only on the output string. Capture is best-effort: if it is unavailable,sessionisNoneand the experiment runs normally.
- Golden Datasets (New): Build evaluation datasets from real production traffic by promoting spans into a dataset.
- Enhancements
fiddler-otelis now a base dependency: Required by trace capture, so it installs with the SDK rather than as an extra.- Pre-production application guard: If an existing
FiddlerClientis scoped to a different application than the dataset’s,evaluate()raisesValueErrorinstead of sending evaluation traces to that application.
- Fixes
evaluate()returns an up-to-date experiment:ExperimentResult.experimentreflectedcreate()-time state, sostatusreadPENDINGandduration_mswasNoneeven after a completed run. The runner now reassigns the hydrated entity at each status transition.
0.5
- Fixes
- Experiment runner threading: Fixed an issue where the experiment runner could fail when executed inside a thread.
0.4
- Breaking Changes
- Experiment result models renamed: The experiment item data models were refactored.
ExperimentItemis nowExperimentResultItemandExperimentItemResultis nowScoredExperimentItem(with a newExperimentResultScoremodel).Experiment.get_items()now returnsIterator[ExperimentResultItem]. Update imports fromfiddler_evals.pydantic_models.experimentaccordingly.
- Experiment result models renamed: The experiment item data models were refactored.
- Enhancements
- CustomJudge alignment:
CustomJudgenow aligns with Fiddler’s internal PromptSpecV2 model for consistent custom LLM-as-a-Judge evaluation.
- CustomJudge alignment:
- Fixes
- Fixed a URL construction issue affecting API requests.
0.3
- New Evaluators
- Context Relevance (New): Measures whether retrieved documents are relevant to the user query. Ordinal scoring — High (1.0), Medium (0.5), Low (0.0) with detailed reasoning.
- RAG Faithfulness (New): LLM-as-a-Judge evaluator that assesses whether the response is grounded in the retrieved documents. Binary scoring — Yes (1.0) / No (0.0) with detailed reasoning.
- CustomJudge (New): Build custom LLM-as-a-Judge evaluators using
prompt_templatewith Jinja{{ placeholder }}syntax andoutput_fieldsfor structured evaluation results.
- Enhancements
- Answer Relevance 2.0: Upgraded from binary to ordinal scoring — High (1.0), Medium (0.5), Low (0.0) with detailed reasoning.
- Ordinal Score Bounding: Ordinal scores from the scoring API are now bounded to [0, 1].
- Enhancements
- Model and Credential Parameters:
modelandcredentialare now parameters on LLM-as-a-Judge evaluators, enabling configuration of the LLM used for evaluation. - Evaluator-Level Score Function Mapping: Evaluators now support
score_fn_kwargs_mappingat the evaluator level for more flexible parameter binding. - Score Name Prefix: Added support for custom score name prefixes on evaluators.
- Evals API Error Handling: Improved error handling and messaging for Evals API responses.
- Coherence Prompt Input Required: The
promptinput for the Coherence evaluator is now required. - Removed Pandas Core Dependency: Pandas moved from core to optional dependency, reducing install footprint.
- Docstring Standardization: Fixed docstring errors and standardized documentation format across all evaluators.
- Model and Credential Parameters:
- Removals
- Toxicity Evaluator Removed: The Toxicity evaluator has been removed from the SDK.
- Initial Release
- Core SDK with HTTP client, entity management (Project, Application, Dataset, Experiment), and the
evaluate()function for running experiments. - Evaluators: AnswerRelevance, Coherence, Conciseness, Sentiment, TopicClassification, FTLPromptSafety, FTLResponseFaithfulness, RegexSearch, and support for user-defined function evaluators.
- Data Input: Load test cases from pandas DataFrames, CSV files, or JSONL files.
- Concurrent Processing: Parallel evaluation with ThreadPoolExecutor and tqdm progress tracking.
- PyPI Publishing: Available as
pip install fiddler-evals.
- Core SDK with HTTP client, entity management (Project, Application, Dataset, Experiment), and the