Preloader

Distributed system groups English text more accurately than conventional methods, tests find

0

Distributed system groups English text more accurately than conventional methods, tests find

Lisa Lock

Scientific Editor

Andrew Zinin

Chief Editor

Spark flies to organize text
Spark architecture. Credit: International Journal of Intelligent Information and Database Systems (2026). DOI: 10.1504/ijiids.2026.156312

A distributed computing framework can group large volumes of English text more accurately and efficiently than older methods, according to research published in the International Journal of Intelligent Information and Database Systems. The system could support faster analysis in education, government and business by automatically sorting related blocks of text into groups, or clusters.

The work addresses a practical problem in digital transformation: Much of the information generated by organizations is unstructured text whose meaning can vary with context. Conventional clustering methods are often slow or produce inconsistent groupings when handling large, high-dimensional data sets.

The researchers used Apache Spark, a platform for processing data across multiple computers, with an improved K-means algorithm. K-means groups data according to similarities. The researchers modified the method to use density peaks and maximum-minimum criteria to select better starting points for clusters, rather than relying on random initialization.

In tests on multiple data sets, the new approach improved clustering accuracy by more than 10% compared with conventional K-means methods. The system remained stable as the researchers increased the volume of data processed. They also improved parallel performance by adding computing nodes.

The framework could help organizations organize and extract information from growing collections of English-language documents more efficiently while providing a more consistent basis for subsequent data analysis and decision-making.

Publication details

Xiaochao Yao, English digital transformation algorithm for distributed big data based on Spark, International Journal of Intelligent Information and Database Systems (2026). DOI: 10.1504/ijiids.2026.156312

Provided by
Inderscience

Who’s behind this story?

Lisa Lock

Lisa Lock

BA art history, MA material culture. Former museum editor, paramedic, and transplant coordinator. Editing for Science X since 2021.

Full profile →


Andrew Zinin

Andrew Zinin

Master’s in physics with research experience. Long-time science news enthusiast. Plays key role in Science X’s editorial success.

Full profile →

Citation:
Distributed system groups English text more accurately than conventional methods, tests find (2026, September 23)
retrieved 24 September 2026
from https://techxplore.com/news/2026-09-groups-english-text-accurately-conventional.html
This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no
part may be reproduced without the written permission. The content is provided for information purposes only.

Source: Tech Xplore

Choose your Reaction!
Leave a Comment