More configuration options for video experiments ๐ฅ
0.11.12
Survey experiments involving videos now have additional configuration options for controlling the size of videos, while ACR video experiments now also support references.
Stay up to date with the latest features and improvements of the Mabyduck subjective testing platform.
0.11.12
Survey experiments involving videos now have additional configuration options for controlling the size of videos, while ACR video experiments now also support references.
0.11.11
We extended our support for markers, allowing you to highlight time ranges in audio and videos for raters to focus on.
You can now also adjust the compensation paid to raters for each job, allowing for faster completion times of high-priority jobs.
0.11.10
We added a new "exhaustive" sampling strategy for pairwise experiments. This strategy will sample every possible pair of a dataset in pseudorandom order.
We also replaced our documentation backend, which will make it easier for us to keep documentation up-to-date and in sync with features in our app and API.
0.11.9
Radar charts are now available in the results dashboard for experiments evaluating content along multiple dimensions.
We also improved support for long videos by adding time indicators. For example, the UI of video experiments will now show the remaining watch time for experiments with "minimum play durations" exceeding 3 seconds.
0.11.8
We've added configuration options to give you more control over the size of videos in pairwise video experiments. For example, references can now be scaled down relative to the other videos.
0.11.7
Experiments can now be tagged, which makes it easier to group and organize them.
Custom rater pools now indicate how many raters are contained in them.
Finally, we introduced organization-level API keys and a new endpoint to fetch billing events of organization billing accounts.
0.11.6
We added self-managed custom rater pools, giving you full control over the demographics of a rater pool.
We also made search bars all across the app a lot more powerful.
0.11.4
Our API now lets you retrieve a project's billing events.
We also added a new "Top-1" sampling strategy which automatically focuses resources to identify the best-performing method.
0.11.3
SSO via Okta is now available for our enterprise customers.
We introduced session health scores, which indicate if raters displayed any unexpected behavior.
We also added support for project-level spending limits for organizations.
0.11.2
We improved the processing speed for large datasets of externally hosted files.
0.11.1
We added support for studies in Arabic.
0.11.0
Device requirements checked at the beginning of a session now take into account if datasets contain any 10-bit videos.
0.10.2
We updated our workflows API to be able to handle large-scale custom strategies.
0.10.1
Pairwise audio experiments now support evaluation of audio along multiple criteria, matching our pairwise video experiments.
0.10.0
This release contained several small improvements:
0.9.14
We added support for explicit golden datasets and golden questions mixed into a session.
0.9.13
Project and dataset secrets now support an allowlist to restrict the endpoints with which secrets can be shared.
0.9.12
This release comes with a variety of new features.
0.9.11
Pairwise video experiments now also have sequential video playback as an option.
0.9.10
We improved the robustness of our AI rater pools.
This release brings updates to our session and rater health scores which we use to evaluate the performance of raters.
0.9.9
You can now set up project-wide secrets which can be used across embedded experiments.
We've also made it possible to filter plots by parameters, and embedded experiments can now store custom metadata on slates.
0.9.8
Audio surveys now support a new "highlight transcript" question type, allowing raters to highlight words in a transcript. Survey experiments also gain a new "Percent chosen" metric for radio buttons and checkboxes, with new plots of these metrics on the results pages.
0.9.7
We improved the calibration of Elo score error bars, which also improves our active sampling strategies.
0.9.6
This release comes with a variety of improvements.
0.9.5
We added automatic checks of video encoding settings for self-hosted videos. For example, to check if videos have been fragmented.
0.9.4
We updated API keys to be associated with users.
0.9.3
You can now set up project webhooks for dataset status changes, job completion, and session completion.
0.9.2
Hero videos included in introductions can now be configured to take over the full screen. They are also now configurable directly via the UI, not just via the API.
0.9.1
Pairwise video experiments can now include arbitrary checkboxes.
We've also added Elo support for MUSHRA experiments.
0.9.0
Experiment introductions can now include a custom hero image, configurable via the API.
0.8.8
This release contains a couple of new features:
0.8.7
This release comes with a variety of improvements.
0.8.6
We introduced Elo scores for ACR experiments. These scores are based on a Plackett-Luce model which ignores the absolute scores but only considers how an individual rater would rank conditions.
0.8.5
We added new response types and support for multiple dimensions in pairwise video experiments.
It is now possible to create datasets, experiments, and jobs in a single API request via our new workflows API. This endpoint also supports specifying a fully custom strategy.
This release also added custom S3 storage support for our enterprise customers.
0.8.4
Added a "scale to fit" option for ACR image experiments.
0.8.3
It is now possible to use self-hosted datasets by providing a list of URLs, instead of uploading media files directly to us.
We also improved our API, and it is now possible to launch jobs where previously it was only possible to configure drafts via the API.
0.8.1
Our embedded experiments are now widely available. This type of experiment uses JavaScript to include arbitrary content in experiments, and is ideal for running interactive studies.
0.7.9
We released new types of experiments that allow the configuration of arbitrary surveys below images, audio, or video.
0.7.8
We improved our support for datasets with very large numbers of conditions. This is useful, for example, when you want to collect labels for training and need to label a large number of audio, images, or videos that are not AI-generated.
0.7.7
Selection strategies have received more configuration options. For example, it is now possible to evaluate only a subset of a dataset. It is also possible to always include one method in pairwise comparisons against other methods.
This release also makes it possible to scale (instead of cropping) images in pairwise image experiments.
0.7.5
We added the ability to add configurable plots to rubrics and leaderboards.
It is now also possible to create draft experiments and jobs via our API.
0.7.4
Today, we are opening up Mabyduck to everyone.
0.7.3
We made small tweaks to our design and changes to our backend to prepare for a public launch.
0.7.2
This release contained several improvements:
0.7.1
We added optional confidence regions to line graph visualizations of your results.
This release also adds support for references in ACR audio experiments.
0.7.0
This release comes with a variety of improvements.
0.6.3
We improved the handling of very large dataset uploads through the browser. If a dataset upload is interrupted for any reason, it is now possible to resume uploads.
This version also adds a new Markdown input field for writing introductions.
0.6.2
We introduced the ability for raters to leave feedback on individual slates and alert us to any potential issues with an experiment.
This version also updated the leaderboards' design.
0.2.6
We internationalized our experiments. In addition to English, we now support French and German.
Additionally, different experiments can now use different config files. This allows you to upload a single dataset with multiple config files for different experiments.
0.2.5
Pairwise image experiments now support references. We also introduced new configuration options for the MUSHRA experiment.
0.2.2
We introduced a new pairwise image experiment. We also added a way to preview images in datasets.
0.2.0
We have implemented our own pre-screening protocols. This allows us to provide you with a higher quality of raters whose ability and hardware enable them to detect fine differences between stimuli.
0.1.6
It is now possible to launch experiments to crowd-sourced raters through our platform.
0.1.5
We added configuration options to change how waveforms are rendered in MUSHRA experiments. In particular, it is now possible to only render the waveform of the reference so that raters can not draw conclusions based on the waveform.
0.1.4
Release 0.1.4 is packed with new features:
0.1.3
We addressed some minor bugs in the MUSHRA experiment.
0.1.2
Datasets now support config files. These can be used to change the interface for each slate. For example, to display text prompts next to stimuli.
0.1.0
Today, we are excited to release a private beta version of Mabyduck to our design partners.