DETEXA: declarative extensible text exploration and analysis through SQL

Επιστημονική δημοσίευση - Άρθρο Περιοδικού uoadl:3340659 30 Αναγνώσεις

Μονάδα:
Ερευνητικό υλικό ΕΚΠΑ
Τίτλος:
DETEXA: declarative extensible text exploration and analysis through SQL
Γλώσσες Τεκμηρίου:
Αγγλικά
Περίληψη:
Metadata enrichment through text mining techniques is becoming one of the most significant tasks in digital libraries. Due to the exponential increase of open access publications, several new challenges have emerged. Raw data are usually big, unstructured, and come from heterogeneous data sources. In this paper, we introduce a text analysis framework implemented in extended SQL that exploits the scalability characteristics of modern database management systems. The purpose of this framework is to provide the opportunity to build performant end-to-end text mining pipelines which include data harvesting, cleaning, processing, and text analysis at once. SQL is selected due to its declarative nature which offers fast experimentation and the ability to build APIs so that domain experts can edit text mining workflows via easy-to-use graphical interfaces. Our experimental analysis demonstrates that the proposed framework is very effective and achieves significant speedup, up to three times faster, in common use cases compared to other popular approaches. © 2023, The Author(s).
Έτος δημοσίευσης:
2023
Συγγραφείς:
Foufoulas, Y.
Zacharia, E.
Dimitropoulos, H.
Manola, N.
Ioannidis, Y.
Περιοδικό:
International Journal on Digital Libraries
Εκδότης:
Springer Science and Business Media Deutschland GmbH
Επίσημο URL (Εκδότης):
DOI:
10.1007/s00799-023-00358-1
Το ψηφιακό υλικό του τεκμηρίου δεν είναι διαθέσιμο.