7 things AI needs from your test data, and how Simcenter Testlab brings a complete AI-ready data workflow
What’s new in Simcenter Testlab 2606
Simcenter Testlab 2606 introduces a complete AI-ready data workflow that helps engineers acquire, validate, label, manage, and deliver test data for AI and machine learning applications. By connecting data acquisition, traceability, workflow automation, and data management in a single environment, organizations can turn trusted test data into AI-ready assets with less manual preparation and greater confidence.
The promise of AI in testing and validation is easy to state: use the data you already have to predict performance earlier, build fewer physical prototypes, and shorten the design cycle. Once a model is trained, it will predict in seconds what a test campaign would take weeks to measure. With products getting more complex and timelines getting shorter, nobody needs convincing that this is worth a try.
But this is where the difficulties begin. An AI model is only as good as the data it learned from, and historical test data was rarely captured with future training in mind. We used to test to answer a specific engineering question, but now we have to make sure that our test data is not only correct, but also AI-ready. Otherwise, before a first model can be trained, test experts and data scientists will spend a significant amount of their time preparing and cleaning the data instead.

So before you dive deep into the world of machine learning, check if you can answer the following seven questions:
| Quality | Has your data been captured with high-quality data acquisition hardware, that you can trust? |
| Validity | Has someone confirmed these results are correct, or are they simply the files that were saved? |
| Traceability | Can you tell what result has been measured, on which configuration, under what conditions, without asking the engineer who ran the test? |
| Completeness | Do your campaigns cover the operating conditions the model will be asked to predict, including the edge cases? |
| Quantity | Do you have enough data to train on, and a way to generate more if you don’t? |
| Accessibility | Can your data scientist reach the results without requesting a file by email or a Teams message? |
| Connectivity | Can you automate the complete workflow or will you have to do this all by hand? |
If you answered “no” more than once, do not be concerned. That gap is exactly what we set out to close with the 2606 release of Simcenter Testlab.
In the sections that follow I’ll address each of the seven challenges and help you start delivering AI-ready data with every test result.
Quality – are your measurements actually trustworthy?
An AI model cannot tell the difference between a physical effect and a sensor fault. It will simply learn both. A loose accelerometer, a clipped channel, a microphone that reads too low for the whole test will all affect your data in a negative way if you cannot react on the spot.
This is why quality starts at the front end. Simcenter SCADAS is built to keep signal integrity at the point of measurement, also with high channel counts and in harsh test environments, and it stores the overload, connectivity and calibration information together with the data instead of in a separate logbook.
If you have already decided to choose Simcenter SCADAS as your data acquisition hardware of choice, the quality requirement is largely handled. If you’re still looking for your perfect DAQ, check this great post about Simcenter SCADAS RS.

Validity – are your results correct?
Validating test data is an thankless task, which can consume hours of your time if done manually. Luckily, Simcenter Testlab Process Designer has plenty of methods designed especially to automate time data, block result and single values validation. Below is a simple example

This approach is not only limited to time data validation. You can easily check block data against specific levels or target curves to further automate your validation. Learn more about data validation in Simcenter Testlab Process Designer by watching this free online webinar.
Traceability – what have you actually measured?
Having hundreds of measured results ready to be used as input to a machine learning workflow means nothing, if you cannot distinguish between measurement A and B. Correctly applied labels are a key requirement to build KPI prediction models. Luckily, Simcenter Testlab comes with a built-in Simcenter Descriptive Data Model And Template Editor – a free of charge tool to design an ASAM compliant annotation model that fits your data.

You can fill-in 80-90% of your annotation even before a measurement campaign is even started, leaving only a limited number of fields that need to be completed on the spot, such as the weather condition information, to the colleagues performing the tests.
Check this “old but gold” blog from Dries to learn more about annotating your data: Data is silver, data & metadata is gold.
Completeness – have you tested all the conditions?
A reliable model will need to be trained with consistent and complete datasets. For example, if you plan to predict road noise NVH performance, all vehicles should be tested on a similar type of maneuver.

In 2606 release, we’ve added a new workbook to the Simcenter Testlab Neo family, the Schedule Acquisition. You can now freely design and automate complete measurement campaigns, from instrumentation, operator instructions, list of measurement tasks, annotation, processing up to reporting.

Scheduling enforces the same task list, settings and conditions on every test, allowing results to stay comparable, campaign to campaign. It’s also a great way to ensure that operators are guided through every step of the process, lowering the skill barrier required to execute a test procedure designed by an expert.

For more advanced in-house procedures, you can also include user-defined actions that will trigger external applications, such as executables or Python scripts.

Not planning to measure with a PC? No worries, you can always use a schedule on your SCADAS RS.
Quantity – do you have enough data?
Test engineers are continuously pressured to reduce the number of physical prototypes and therefore conduct less tests. But this contradicts the requirement of having sufficient data to be used as input to an AI model training. In those situations, Simcenter Testlab Neo can help complement measured data with siumulated data.
Using technologies such as component-based Transfer Path Analysis (C-TPA), and virtual prototype assembly, you can extend your datasets and help close data gaps with data you can trust.

This creates a stronger foundation for downstream AI applications while making better use of both physical and simulation engineering data.
Accessibility – can you find your data?
Even if you’ve taken all the required measurements, extracted all the relevant KPIs and added all the required annotation, you might still struggle with building machine learning workflows if you simply cannot find your measurements.

This situation can be easily avoided by investing in solutions such as the Simcenter Testlab Data Management, a centrally deployed ASAM-ODS database to store and search your test data.

Connectivity – can you automate the tedious tasks?
The glue holding these requirements together is the ability to automate the data validation, KPI extraction, labelling, publishing and retrieval of data before connecting your test data to AI platforms, such as Rapidminer AI Studio.

Simcenter Testlab Workflow Automation (TWA) has been extended in the 2606 release with the option to search and process data stored directly in Simcenter Testlab Data Management. Simply define a search, such as “find me all C-segment hybrid vehicle measurements taken on proving ground in Spain”, sit back and enjoy your coffee, as TWA will take care of retrieving all your data, re-processing it as needed and exporting it towards your AI training platform of choice.

Bringing the whole workflow together
With Simcenter Testlab Neo 2606, we bring the complete next-generation workflow of acquiring, labeling, processing and delivering of AI-ready data together in one integrated environment.
