A corpus-based study of idiomatic expressions: Focusing on the resolution of ambiguity
Network analysis surveys corpus-based idiom studies and shows contextual embeddings outperform word-based models for ambiguity resolution and Korean idiom classification.
Published May 30, 20221 citationPaper ↗
69%
OverallHighly rated
?
OverallHighly ratedVote to see the score
Readers
–
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel3/20reviewers recommend it
lenient 2/5
medium 1/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Abstract
이 논문의 목적은 관용어 표현에 대한 말뭉치 기반 연구의 다양한 분야를 네트워크 분석을 통해 확인하고, 자연어 처리 분야에서 관용어 표현 연구의 결과를 소개하는 데 있다. 실제 언어 사용 데이터인 말뭉치 기반 연구를 통해 관용 표현 연구의 세부 주제가 확장되고 응용 분야와의 접목이 활발해졌다. 언어 처리 분야의 관용 표현의 중의성 해소연구는 딥러닝의 등장으로 다시 활성화되기 시작하였으며 관용 표현 주석 말뭉치가 중요한 역할을 수행하고 있다. 관용 표현과 일반 표현의 중의성 해소를 위해 어절 단위, 형태소 단위 말뭉치에 대해 다양한 층위의 임베딩을 생성하여 사전 학습 임베딩 및 분류 태스크 실험을 수행한 결과 자동 분류의 가능성을 확인하였다. 한국어 관용 표현 탐지와 분류 실험을 통해 문맥이 반영된 임베딩의 성능이 단어 기반 임베딩보다 우수하다는 결과를 도출하였다.