Tether Launches Extensive AI Training Dataset to Enhance Local Model Capabilities
Tether AI Research has unveiled QVAC Genesis III, a synthetic training dataset comprising 191.4 billion tokens derived from 159.6 million documents. This initiative aims to bolster the reasoning and teaching capabilities of smaller artificial intelligence models, enabling them to operate efficiently on personal devices such as laptops and mobile phones, thereby reducing dependence on large cloud-based models.
The dataset spans 19 curriculum-related fields, including biology, chemistry, physics, mathematics, computer science, medicine, electrical engineering, and machine learning, targeting applications for high school, university, and professional education. Tether emphasizes that the focus of Genesis III is to train models to elucidate problem-solving processes, identify flawed reasoning, and offer corrections, rather than merely providing final answers.
In preliminary tests, a model with 1.7 billion parameters trained using the Genesis III dataset achieved an impressive effective answer rate of 99.45% on relevant benchmarks. When compared to the Cosmopedia-v2 training model of similar size, Genesis III demonstrated significant improvements in scores across various STEM benchmarks, including increases of 28.57, 21.35, and 15.03 percentage points on the ARC-Easy, ARC-Challenge, and MMLU assessments, respectively.
Source: KLEA News