Essentia: Mining Domain-Specific Paraphrases with Word-Alignment Graphs

Ma, Danni; Chen, Chen; Golshan, Behzad; Tan, Wang-Chiew

Computer Science > Computation and Language

arXiv:1910.00637 (cs)

[Submitted on 1 Oct 2019 (v1), last revised 4 Oct 2019 (this version, v2)]

Title:Essentia: Mining Domain-Specific Paraphrases with Word-Alignment Graphs

Authors:Danni Ma, Chen Chen, Behzad Golshan, Wang-Chiew Tan

View PDF

Abstract:Paraphrases are important linguistic resources for a wide variety of NLP applications. Many techniques for automatic paraphrase mining from general corpora have been proposed. While these techniques are successful at discovering generic paraphrases, they often fail to identify domain-specific paraphrases (e.g., {staff, concierge} in the hospitality domain). This is because current techniques are often based on statistical methods, while domain-specific corpora are too small to fit statistical methods. In this paper, we present an unsupervised graph-based technique to mine paraphrases from a small set of sentences that roughly share the same topic or intent. Our system, Essentia, relies on word-alignment techniques to create a word-alignment graph that merges and organizes tokens from input sentences. The resulting graph is then used to generate candidate paraphrases. We demonstrate that our system obtains high-quality paraphrases, as evaluated by crowd workers. We further show that the majority of the identified paraphrases are domain-specific and thus complement existing paraphrase databases.

Comments:	accepted at the 13th Workshop on Graph-Based Natural Language Processing
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:1910.00637 [cs.CL]
	(or arXiv:1910.00637v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1910.00637

Submission history

From: Danni Ma [view email]
[v1] Tue, 1 Oct 2019 19:51:57 UTC (144 KB)
[v2] Fri, 4 Oct 2019 19:04:53 UTC (144 KB)

Computer Science > Computation and Language

Title:Essentia: Mining Domain-Specific Paraphrases with Word-Alignment Graphs

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Essentia: Mining Domain-Specific Paraphrases with Word-Alignment Graphs

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators