Maximizing fusion performance in tokamaks relies on high energy confinement, often achieved through distinct operating regimes. The automated labeling of these confinement states is crucial to enable large-scale analyses or for real-time control applications. While this task becomes difficult to automate near state transitions or in marginal scenarios, much success has been achieved with data-driven models. However, these methods generally provide predictions as point estimates, and cannot adequately deal with missing and/or broken input signals. To enable wide-range applicability, we develop methods for confinement state classification with uncertainty quantification and model robustness. We focus on off-line analysis for TCV discharges, distinguishing L-mode, H-mode, and an in-between dithering phase (D). We propose ensembling data-driven methods on two axes: model formulations and feature sets. The former considers a dynamic formulation based on a recurrent Fourier neural operator-architecture and a static formulation based on gradient-boosted decision trees. These models are trained using multiple feature groupings categorized by diagnostic system or physical quantity. A dataset of 302 TCV discharges is fully labeled, and we release it publicly to encourage the community to build upon this work. We evaluate our method quantitatively using Cohen’s kappa coefficient for predictive performance and the expected calibration error for the uncertainty calibration. Furthermore, we discuss performance using a variety of common and alternative scenarios, the performance of individual components, out-of-distribution performance, cases of broken or missing signals, and evaluate conditionally-averaged behavior around different state transitions. Overall, the proposed method can distinguish L, D and H-mode with high performance, can cope with missing or broken signals, and provides meaningful uncertainty estimates.
トカマクにおける核融合性能の最大化は、高いエネルギー閉じ込めに依存しており、それはしばしば明確な運転領域を通じて達成される。これらの閉じ込め状態の自動ラベル付けは、大規模解析やリアルタイム制御アプリケーションを可能にするために極めて重要である。このタスクは状態遷移の近傍や限界的なシナリオでは自動化が困難になるが、データ駆動モデルでは多くの成功が達成されている。しかしながら、これらの手法は一般に点推定として予測を提供し、欠落した、および/または破損した入力信号に適切に対処できない。広範囲での適用可能性を可能にするために、我々は不確実性定量化とモデル頑健性を備えた閉じ込め状態分類の手法を開発する。我々はTCVディスチャージのオフライン解析に焦点を当て、Lモード、Hモード、および中間のディザリング位相(D)を区別する。我々はデータ駆動手法を2つの軸、すなわちモデル定式化と特徴量セットでアンサンブルすることを提案する。前者は、リカレントFourierニューラルオペレーターアーキテクチャに基づく動的定式化と、勾配ブースティング決定木に基づく静的定式化を考慮する。これらのモデルは、診断システムまたは物理量によって分類された複数の特徴量グループを用いて訓練される。302のTCVディスチャージからなるデータセットが完全にラベル付けされており、我々はこれを公開し、コミュニティがこの研究を基に構築することを促進する。我々は、予測性能にはCohenのκ係数を、不確実性較正には期待較正誤差を用いて、定量的に本手法を評価する。さらに、多様な一般的および代替シナリオを用いた性能、個々の構成要素の性能、分布外性能、破損または欠落した信号のケースについて議論し、異なる状態遷移周辺の条件付き平均挙動を評価する。全体として、提案手法はL、D、Hモードを高い性能で区別でき、欠落または破損した信号に対処でき、有意義な不確実性推定を提供する。