As new programs for AI-generated images, text, and videos have become available over the past few years, concerns over the accuracy and quality of the generated information have become apparent. How can we measure these results to determine their accuracy?
AI programs have been widely used by the public to create text, produce images of unique scenarios, and now develop videos of different scenes, all at the hands of the “prompt” for generation. With new programs being developed and released to the public, some inaccuracies have occurred in the resulting information generated, further pushing the correction and advancement of these programs to become more accurate.
One of the newest programs set to be released to the public some time in 2024 is OpenAI’s Sora. According to OpenAI, “Sora is an AI model that can create realistic and imaginative scenes from text instructions,” generating “videos up to a minute long.” As this new program is soon to be available to the public, concerns over the accuracy of the results in terms of looking like reality have created criteria for the ways in which we measure the accuracies and overall quality of the results of AI-generated information.
There are four measurements that we can use to determine the accuracy of results from programs like Sora, including measurements of quality, temporal coherence, if we received what we asked for, and of the result maintained object permanence and consistency.
- Quality
This first measurement of the quality of the results determines how realistic the images and videos are. When looking at a result of a video or image of a person, realistic qualities could be determined by looking at if the results include an entire person or if some body parts are missing or look warped, or if the results look cartoon-like instead of realistic. By determining the quality of the results, these images and videos can become very convincing and look as if they naturally occurred rather than being AI-generated.
- Temporal Coherence
This second measurement of temporal coherence determines if the results make sense in sequence and in time. This element of the AI-generated result determines how smooth one moment of a visual transitions to the next. When looking at Sora’s visual generation, we can see how smooth the overall video is when moving through a scene, as opposed to a more stop-motion-looking video that is not as smooth in transitioning between scenes. By determining the temporal coherence quality of a visual, an accurate AI-generated video can look as if it was taken with a camera, creating a clear image of a potentially realistic scene.
- Prompt Results
The third measurement of accuracy of results is by seeing if the results are aligned with what the prompt stated. In other words, did we get what we asked for in the prompt. When typing a prompt into an AI-generation program, sometimes the results are slightly or entirely far off from the request, creating inaccuracies in the results and causing the need for additional prompts to be used. If results are accurate, then the correlation between what the prompt is and what the results are will be high.
- Object Permanence and Consistency
The fourth measurement of accuracy is in determining object permanency and consistency in results. When generating a video in AI-programs with objects moving in front of each other or the view moving around a scene, accurate results will show objects in the correct places before and after they are in the frame. For example, in a generated scene where a person walks in front of a sign, an accurate result will show the exact same sign from before they walked in front of it to after. By measuring the object permanence and consistency of a visual, the results can be accurate in showing a realistic scene as it would be in real life.
When determining the accuracy of AI-generated information, these measurements can be used to test the reliability of the results and if they are realistic. Even though these programs have created more realistic results over time, these programs still have some issues in producing visuals that the programs may not have enough information on. Some programs have “hallucinations” and show zero-shot performance/emergent behavior results where results show a visual or text that did not exist in the training data, or shows something that it was not assigned to do. These results are inaccurate and unrealistic, providing another reason to measure the accuracy of AI-generated imagery.
Each of these measurements determines the overall quality and accuracy of AI-generated visuals. As these programs develop, their visuals have become more accurate and difficult for users to determine what is real and what is AI-generated. By using these measurements and criteria in determining the accuracy and quality of results, users can become more familiar in the signs of what is authentic and what is generated by this new technology.
Sources:
https://openai.com/sora