Wednesday 26 March 2025
The quest for high-quality labeled training data has long been a thorn in the side of machine learning researchers and practitioners alike. The problem is twofold: first, collecting large amounts of accurate labeling requires significant human effort, which can be both time-consuming and costly. Second, even when labels are obtained, they may not accurately reflect the complexities of real-world data.
To address these challenges, a team of researchers has developed ScriptoriumWS, a system that leverages code-generation models to automatically produce programming assistance for synthesizing labeling functions (LFs). The idea is simple: by providing a prompt or set of rules to guide the code generation process, LFs can be created quickly and efficiently, without requiring extensive human expertise.
To evaluate the effectiveness of ScriptoriumWS, the researchers conducted a comprehensive benchmarking exercise using the WRENCH weak supervision dataset. This dataset consists of 101 labeling tasks, each with its own set of rules and constraints, making it an ideal testing ground for any system seeking to automate LFs.
The results were promising: when compared to human-designed LFs, ScriptoriumWS-generated LFs showed comparable accuracy in most cases, while also achieving higher coverage rates. This means that the generated LFs were able to accurately label more data points than their human-crafted counterparts.
But what makes ScriptoriumWS particularly noteworthy is its flexibility and adaptability. By using different prompting strategies, the system can be fine-tuned to produce LFs tailored to specific tasks or domains. For example, providing a set of heuristics or rules can guide the code generation process towards more accurate labeling outcomes.
The implications of ScriptoriumWS are far-reaching: by reducing the time and cost associated with collecting high-quality labeled data, researchers and practitioners can focus on developing more sophisticated machine learning models that better serve real-world applications. Moreover, the system’s ability to adapt to different domains and tasks opens up new possibilities for weak supervision in areas such as natural language processing, computer vision, and beyond.
Of course, like any system, ScriptoriumWS is not without its limitations. The quality of the generated LFs will ultimately depend on the quality of the prompting strategy used, as well as the capabilities of the code-generation model itself. Furthermore, the system’s performance may degrade in scenarios where the rules or heuristics provided are incomplete or inaccurate.
Cite this article: “Automating Labeling Functions with ScriptoriumWS”, The Science Archive, 2025.
Machine Learning, Data Labeling, Code Generation, Weak Supervision, Labeling Functions, Automated Lfs, Programming Assistance, Wrench Dataset, Benchmarking Exercise, High-Quality Training Data







