Evaluation
The purpose of creating a Project and performing several Experiment in its frame is basically to train the best-performing ModelVersion. As detailed on the previous page, you can log Metrics in the Experiment Tracking dashboard to get the insights needed to assess the quality of the training performed. In addition to the Metrics, Picsellia also lets you compare, on a dedicated bench of images, the predictions made by the freshly trained model with the ground truth — this is called an Evaluation.
By leveraging the Evaluation interface, you can get a precise overview of the model's behavior, get performance metrics on each evaluated image, and identify the contexts in which the model performs well or not.
With Metrics and Evaluation, you'll have all the insights needed to assess the quality of the trained ModelVersion, understand why it behaves the way it does, and identify ways to improve its performance in a further Experiment.
1. Create an Evaluation
Evaluation are computed and created by your training script. It needs to get a bunch of images unused during the training phase and already annotated in an existing DatasetVersion, in order to use them as ground truth.
This bunch of images can come from the training script splitting the DatasetVersion attached to the Experiment, or from a DatasetVersion dedicated to Evaluation and attached to the Experiment with an alias indicating its purpose.
The script then makes the just-trained ModelVersion infer on the evaluation images and sends the result to the Picsellia Experiment.
The outcome is that, for each evaluation image, you can compare the prediction made by the ModelVersion with the ground truth pulled from the original DatasetVersion. You'll also have access to evaluation metrics computed by Picsellia based on the comparison of predictions and ground truth.
The creation of Evaluation is already integrated into almost all the training scripts attached to ModelVersion from the Public Registry.
However, if you are using your own training script integrated with Picsellia, you can use this dedicated tutorial, detailing how to create and log Evaluation within your Picsellia Experiment.
2. Evaluation interface
From an Experiment, you can access the Evaluations tab, as shown below:

Access Evaluation tab
After Experiment creation, the Evaluation interface is empty. It is only once the training script has been executed successfully that Evaluation are created in the Evaluation interface.
The Evaluation interface offers a way to visualize and explore Evaluation with the same philosophy and features as the Datalake to browse among data, or a DatasetVersion to browse among Asset.

Evaluation interface
A. Evaluation visualization
a. Evaluation views and global navigation
You'll find three different views:
- Grid view, letting you visualize many
Evaluationat a time. - Table view, letting you quickly visualize many
Evaluationand their associated information (filename, width, height, detection performances). - Details view, letting you visualize one
Evaluationin detail (full-screen image display and associated information).

Switch between views
The Grid, Table, and Details views, along with their navigation and display options, are in substance the same ones already covered for the Datalake (detailed here) and a DatasetVersion (detailed here).
b. Evaluation Shape
For each Evaluation, you can visualize the image, as well as the superposition of the prediction made by the trained ModelVersion and the ground truth inherited from the DatasetVersion.
The ground truth Shape are displayed in green, whereas the predicted Shape are red. For each Shape, the associated Label is displayed along with the confidence score in the case of predictions.

Evaluation visualization
To ensure the smoothest visualization possible, you can select which elements to display on your image.
You can choose to display, on your Evaluation image, only the ground truth, only the predictions, or both. You can also filter the Shape to display by Label or by confidence score (in the case of predictions).

Select the annotation to display: Ground Truth, Evaluation or Both

Filter the Evaluation Shapes to display based on the confidence score
c. Evaluation metrics
As explained previously, for each Evaluation, Picsellia compares ground truth and prediction to compute different metrics, allowing you to assess the quality of the prediction made by the ModelVersion on that particular image.
On the Details view, besides the usual properties, Metadata, and Custom Metadata, some metrics computed by the platform are displayed to assess the performance of the model against the ground truth for each Evaluation.
The number of ground truth Shape, True Positives, False Positives, and False Negatives are displayed as shown below:

TP, FP, FN scores
In addition, more in-depth metrics (AP/AR on several IoUs) can be visualized in the Table and Details views for each Evaluation:

AP/AR on several IoUs
Not applicable for ClassificationThose metrics are computed only for
Evaluationperformed in the frame of anExperimentwith Inference Type Object Detection or Segmentation. For Classification, please refer to the dedicated section at the end of this page.
B. Evaluation filtering
To let you investigate the Evaluation generated and understand the ModelVersion's behavior in depth, Evaluation can be filtered on several criteria.
a. Date
The date picker lets you filter Evaluation on their creation date by clicking on the Select Date button. The date picker will open, letting you select the timeframe to search on.

Filter on Evaluation creation date
Once the date range is selected, the associated query is filled in the search bar, letting you complete it directly through the search bar for a more specific query.
b. Metrics
The Evaluation interface also lets you filter on the computed metrics. To do so, click on + Add filter and define, through the modal, the metric and the values to filter on. For instance, let's create a filter that displays only the Evaluation created between 09/19/2023 and 09/22/2023, with an Average Precision (IoU 50-95) between 0.291 and 0.782:

Filter on Evaluation metrics
The images matching the active filters will be displayed.
Any filter can be removed at any time by clicking on it, then on the trash icon.
c. Custom Metadata
The + Add filter modal also lets you filter Evaluation on the Custom Metadata of the underlying Data — the same key/value fields you can define on a Data in the Datalake. Pick the key you want to filter on — the keys declared in the Custom Metadata schema of the Datalake the underlying Data belongs to are suggested, detailed here — then its value type, and the value or range to match. This is useful to slice your evaluation bench along whatever Custom Metadata you already track on your Data, for instance filtering Evaluation down to the ones acquired by a specific sensor or under a specific condition, without needing an Asset Attribute or Shape Attribute for it.

Creation of a filter based on a Custom Metadata value
d. Filter Evaluations by AssetAttribute and/or ShapeAttribute
Evaluations by AssetAttribute and/or ShapeAttributeThe + Add filter modal also lets you filter Evaluation on Attribute values — those inherited from the ground truth Asset, and those carried by its Shape — as well as on the Shape themselves, by their type or their Label. This is a direct way to check the ModelVersion behavior on a specific slice of the evaluation bench, for instance only the images whose Asset Attribute marks a difficult acquisition condition. Attribute filtering is detailed here.

Creation of a filter based on Shapes Attributes values
Summary on filtersIn summary, several types of filters can be combined to browse your
Evaluationbased on multiple criteria — creation date, average precision, average recall,Custom Metadata,Asset Attribute, orShape Attribute— alongside the Query Language (detailed below), to refine your search even further.
C. Search Bar
As with all the image visualization views available on Picsellia, you have access to the search bar powered by our Query Language. This search bar lets you create complex queries leveraging all the properties related to each Evaluation.
The Query Language lets you browse among properties of the image itself (linked to an Asset), of the ground truth Annotation, or of the Evaluation Annotation. You can rely on auto-completion to see all the properties you can search on.

Create complex queries with the Query Language
To search on the Evaluation Annotation, you need to type polygons.xxx (or the current shape type), whereas if you want to filter on the ground truth Annotation, you need to access the Asset first and type asset.polygons.xxx.
D. Evaluation deletion
To delete any computed Evaluation, select the Evaluation to delete and click on Delete.

Evaluation deletion
E. Embeddings and visual search
Embeddings can also be computed for the Evaluation of an Experiment, the same way as for the Data of a Datalake (detailed here). On an evaluation bench of a few thousand images, this is how you find the images that resemble the one the ModelVersion failed on, instead of scrolling until you meet them.
The computation is activated from the Data Embeddings page of the Experiment Settings, where you can also follow its progress, Retry the Evaluation that could not be computed, and Deactivate the computation, which deletes the embeddings already computed.

Embeddings activation from Evaluation Settings page
Once the computation is over, you can look for the Evaluation that look like a selected one, or that match a text prompt.

Exploring Evaluation using similarity search or text-to-image search
You can also switch to the Embeddings exploration mode to visualize a UMAP projection of every Evaluation embedding as a scatter plot, automatically grouped into clusters of visually similar images. Selecting a cluster with the lasso, rectangle, or cluster-selection tool highlights the corresponding Evaluation in the grid — a fast way to check whether a cluster of visually similar images shares the same kind of ground truth/prediction mismatch, instead of inspecting Evaluation one by one.

Exploring Evaluations with UMAP projection
F. Charts
Alongside Query Language and Embeddings, a third exploration mode, Charts, lets you build aggregate visualizations over your Evaluation — for instance, the average precision of your Evaluation, grouped by DataTag, to spot which subsets of your data the ModelVersion struggles with. Creating and managing a Chart is detailed in Charts.

Example of a Chart based on AP_50_95, and the Evaluation associated with a given bar
G. Compare across Experiment
Everything so far looks at the Evaluation of a single Experiment. If two or more Experiment share a common evaluated DatasetVersion, Picsellia also lets you overlay their Evaluation on the same Asset, to see directly, image by image, how their predictions differ. This is detailed in Comparison.
3. Classification
The case of Classification is handled a bit differently than Object Detection or Segmentation in the Evaluation interface. This is mainly because average recall and average precision metrics don't make sense for Classification Experiment.
The creation of Evaluation for Classification by a training script is detailed here.
Most of the features remain the same as detailed previously.
The visualization of ground truth and prediction on each Evaluation follows the same guideline, with the ground truth Label on the left and the predicted Label on the right.

Evaluation of Classification
The other major difference lies in the computed metrics. For Classification Evaluation, it is the confidence score (referred to as Score in the Evaluation interface) of the prediction that is displayed and usable to create filters.
The ground truth and predicted Label are also displayed as fields in the Table and Details views.

Classification labels for Ground Truth vs Evaluation and Confidence score
To quickly access the Evaluation where the ground truth Label differs from the predicted one, click on the Show class errors button, available once the Filters bar has been opened.

Show class errors button
For instance, in the screenshot above, using the Show class errors button shows only the Evaluation where the ground truth Label differs from the predicted one — in this case, 53 Evaluation out of 279 performed.
Updated 4 days ago