Identifying a low-dimensional embedding of a high-dimensional data set allows exploration of the data structure. In this paper we tested some existing manifold learning techniques for discovering such embedding within the multidimensional operational space of a nuclear fusion tokamak. Among the manifold learning methods, the following approaches have been investigated: linear methods, such as principal component analysis and grand tour, and nonlinear methods, such as self-organizing map and its probabilistic variant, generative topographic mapping. In particular, the last two methods allow us to obtain a low-dimensional (typically two-dimensional) map of the high-dimensional operational space of the tokamak.These maps provide a way of visualizing the structure of the high-dimensional plasma parameter space and allow discrimination between regions characterized by a high risk of disruption and those with a low risk of disruption. The data for this study comes from plasma discharges selected from 2005 and up to 2009 at JET. The self-organizing map and generative topographic mapping provide the most benefits in the visualization of very large and high-dimensional datasets. Some measures have been used to evaluate their performance. Special emphasis has been put on the position of outliers and extreme points, map composition, quantization errors and topological errors.
高次元データセットの低次元埋め込みを特定することで、データ構造の探索が可能になる。本論文では、核融合トカマクの多次元運転空間におけるこのような埋め込みを発見するために、既存の多様体学習手法をいくつか検証した。多様体学習手法の中でも、以下のアプローチが調査された:主成分分析やグランドツアーなどの線形手法、および自己組織化マップやその確率的変種である生成的トポグラフィックマッピングなどの非線形手法である。特に、後者の2つの手法により、トカマクの高次元運転空間の低次元(典型的には2次元)マップを得ることができる。これらのマップは、高次元プラズマパラメータ空間の構造を可視化する手段を提供し、破壊リスクが高い領域と低い領域の判別を可能にする。本研究のデータは、JETにおいて2005年から2009年までに選択されたプラズマ放電から得られたものである。自己組織化マップと生成的トポグラフィックマッピングは、非常に大規模で高次元のデータセットの可視化において最も有益である。これらの手法の評価には、いくつかの指標が用いられた。特に、外れ値や極値の位置、マップの構成、量子化誤差、位相幾何学的誤差に重点が置かれた。