Thursday 27 March 2025
The pursuit of creating realistic images from sketches has been a longstanding challenge in the field of computer vision. While significant progress has been made in recent years, there is still much room for improvement. A new approach published in a paper by researchers at the University of Technology Sydney aims to bridge this gap by introducing a novel technique that uses a learnable lightweight mapping network (LCTN) and pre-trained text-to-image diffusion models.
The traditional method of generating images from sketches involves using generative adversarial networks (GANs), which can produce high-quality results but often struggle with capturing the essence of the original sketch. The proposed approach, on the other hand, leverages the strengths of both GANs and diffusion models to create a more accurate and diverse range of images.
The LCTN is a key component in this process, as it enables the mapping of sketches to a latent space that can be manipulated to generate different images. This is achieved by using a sequence of four fully connected hidden layers with batch normalization and ReLU activation functions. The output of the LCTN is then passed through a pre-trained text-to-image diffusion model, which uses stochastic differential equations (SDEs) to generate images from the latent space.
One of the major advantages of this approach is its ability to produce highly realistic images that closely resemble the original sketch. This is achieved by using a combination of spatial and temporal attention mechanisms to focus on specific regions of the image and manipulate them accordingly. Additionally, the use of SDEs allows for the generation of diverse and high-quality images that are not limited by the constraints of traditional GAN-based methods.
The proposed approach has been tested on several benchmark datasets, including Scribble, QMUL, and Flickr20, with impressive results. The authors report significant improvements in image quality and diversity compared to existing state-of-the-art methods, particularly when it comes to generating realistic images from sketches.
While the proposed approach is still in its early stages, it has the potential to revolutionize the field of computer vision by enabling the creation of highly realistic images from sketches. This technology could have a wide range of applications, from artistic rendering and animation to medical imaging and virtual reality.
In addition to its technical merits, this research also highlights the importance of collaboration between academia and industry in advancing the state-of-the-art in computer vision. The authors’ use of pre-trained models and diffusion processes demonstrates the potential for innovative solutions to emerge at the intersection of these two fields.
Cite this article: “Realistic Image Generation from Sketches using Learnable Lightweight Mapping Network”, The Science Archive, 2025.
Computer Vision, Image Generation, Sketch-To-Image, Gans, Diffusion Models, Latent Space, Text-To-Image, Sdes, Attention Mechanisms, Lightweight Mapping Network.







