Ugail

Technology

A New Era of Generative and Multimodal AI

Published on 2/4/2023

Large language models are beginning to change our understanding of what a machine learning system can do. Their importance may extend far beyond language itself. When combined with vision, these models offer a new way of approaching visual computing, where a system may not only recognise what is present in an image but also describe it, relate it to prior knowledge, follow instructions and reason about what it sees. This could move visual computing beyond individual tasks such as classification or recognition towards more general systems capable of working across many problems.

Generative AI adds another important dimension. Instead of relying entirely on expensive collections of real images, synthetic data could be generated to augment training sets, introduce rare examples and explore variations in pose, lighting, environment or appearance that may be difficult to capture experimentally. Used carefully, this could be particularly powerful in areas such as medical imaging, biometrics, engineering and scientific analysis, although ensuring that generated data remains realistic and unbiased will be critical.

Perhaps the most exciting development is the possibility of combining perception, generation and reasoning within the same computational framework. A machine may increasingly be able to look at complex information, draw on large bodies of knowledge, reason about possible explanations and generate useful outputs in response. We may therefore be entering a remarkable period in machine learning, where increasingly general AI systems allow us to tackle large scale problems in ways that few of us could have anticipated only a few years ago.