| Project/Area Number |
22K17957
|
| Research Category |
Grant-in-Aid for Early-Career Scientists
|
| Allocation Type | Multi-year Fund |
| Review Section |
Basic Section 61030:Intelligent informatics-related
|
| Research Institution | Institute of Physical and Chemical Research |
Principal Investigator |
Teranishi Hiroki 国立研究開発法人理化学研究所, 革新知能統合研究センター, 特別研究員 (50899408)
|
| Project Period (FY) |
2022-04-01 – 2025-03-31
|
| Project Status |
Completed (Fiscal Year 2024)
|
| Budget Amount *help |
¥3,380,000 (Direct Cost: ¥2,600,000、Indirect Cost: ¥780,000)
Fiscal Year 2024: ¥1,300,000 (Direct Cost: ¥1,000,000、Indirect Cost: ¥300,000)
Fiscal Year 2023: ¥1,300,000 (Direct Cost: ¥1,000,000、Indirect Cost: ¥300,000)
Fiscal Year 2022: ¥780,000 (Direct Cost: ¥600,000、Indirect Cost: ¥180,000)
|
| Keywords | 構文解析 / 依存構造解析 / 並列構造解析 / 談話構造解析 / 複合語解析 / 文書表現学習 / 文書エンコーディング / 文脈拡張 / 文書検索 / 文脈解析 / 文書エンコーダ / 大規模言語モデル / 自然言語処理 / 依存構造 / 並列構造 / 複文構造 |
| Outline of Research at the Start |
文の構造解析の基盤技術は、単語と単語の間の係り受け関係(依存関係)を明らかにすることであるが、既存技術は近接して現れる単語間の依存関係しか高精度に解析できていない。科学技術論文などの専門的文書には長く複雑な文が頻出し、離れた単語間の依存関係の解析が困難であるため、単語の依存関係に基づくテキストマイニングの性能に影響を及ぼしている。 本研究は、文を長く複雑にする要因となる複文構造や並列構造などの言語現象に着目し、言語現象の性質を利用した単語間の関係解析を試みる。言語現象の構成要素に基づいた解析手法を確立し、長く複雑な文における単語間の依存関係の解析精度向上を目指す。
|
| Outline of Final Research Achievements |
This study aims to improve the accuracy of syntactic parsing for long and complex sentences by investigating the following approaches: chunk segmentation methods based on discourse structures and noun phrases; a head selection approach that enables the model to learn the parsing order; data augmentation for coordinate structure analysis using pretrained models; and document encoding methods that capture broader contextual information. The chunk segmentation approach revealed challenges such as the misidentification of phrase structures and limitations in existing annotations, while the automatic acquisition of parsing order suffered from error propagation and unstable convergence. On the other hand, data augmentation using the T5 model proved effective for identifying coordinate structures in low-resource settings, and the split-and-merge document encoding model demonstrated performance competitive with existing methods.
|
| Academic Significance and Societal Importance of the Research Achievements |
本研究は、従来困難とされてきた長文の構文解析に対し、複数の新規アプローチを総合的に検討し、その限界と可能性を明らかにした点で学術的意義がある。また、事前学習モデルを用いた学習データの生成手法の開発や文書の分割・統合エンコーディング手法の検証については今後の研究への応用も期待される。研究期間を通じて、事前学習モデルの大規模化や生成AIと呼ばれる汎用的なLLMの進展の影響を受け、構文解析の意義や設計について再評価する契機となり、本研究はLLMの推論・思考を補強・拡張するといった構文解析の新たな展開の可能性につながる知見となった。
|