Researchers have built an AI tutoring system for Vietnamese students that keeps all data on local servers instead of shipping it to ChatGPT or another foreign cloud.
The system, called DeepEdu-v1, runs on a new framework named SCALE. It tackles two problems that have stalled local AI tutors: slow response times on consumer GPUs and models that hallucinate when quizzed on Vietnam's national curriculum. SCALE's inference engine groups token retrieval by cluster rather than by chunk, cutting retrieval calls by a factor of 7.7 and trimming initial response latency by about 35 percent. A second layer builds a verified playbook from past tutoring sessions instead of fine-tuning the model, which the researchers say should reduce the system's reliance on English-language training data over time. In its deployed configuration, DeepEdu-v1 ran nearly twice as fast as standard vLLM serving and raised accuracy on complex tasks from 70.0 percent to 79.5 percent, with the biggest gains on financial-reasoning and interactive-agent benchmarks.
The bigger story here is regulatory, not just technical. Vietnam's Decree 53 restricts sending student data abroad, which rules out cloud assistants for any compliant school. That leaves self-hosted open models as the only legal option, and this paper is a direct attempt to make that option fast and accurate enough to run on affordable hardware instead of a data center.
The numbers come from the team's own benchmarks, not classroom deployment at scale, so treat the accuracy jump as a lab result until it survives contact with actual students.