Unmitigated disruptions pose a much more serious threat when large-scale tokamaks are operating in the high performance regime. Machine learning based disruption predictors can exhibit impressive performance. However, their effectiveness is based on a substantial amount of training data. In future reactors, obtaining a substantial amount of disruption data in high performance regimes without risking damage to the machine is highly improbable. Using machine learning to develop disruption predictors on data from the low performance regime and transfer them to the high performance regime is an effective solution for a large reactor-sized tokamak like ITER and beyond. In this study, a number of models are trained using different subsets of data from the HL-2A tokamak experiment. A SHapley Additive exPlanations (SHAP) analysis is executed on the models, revealing that there are different, even contradicting, patterns between different performance regimes. Thus, simply mixing data among different performance regimes will not yield optimal results. Based on this analysis, we propose an instance-based transfer learning technique which trains the model using a dataset generated with an optimized strategy. The strategy involves instance and feature selection based on the physics behind differences in high- and low-performance discharges, as revealed by SHAP model analysis. The TrAdaBoost technique significantly improved the model performance from 0.78 BA (balanced accuracy) to 0.86 BA with a few high-performance operation data.
未緩和ディスラプションは、大型トカマクが高性能領域で運転されている場合、はるかに深刻な脅威となる。機械学習に基づくディスラプション予測器は、印象的な性能を示すことができる。しかし、その有効性は、かなりの量の訓練データに基づいている。将来の核融合炉では、装置への損傷のリスクなしに高性能領域でかなりの量のディスラプションデータを取得することは、極めて困難である。低性能領域のデータを用いて機械学習によりディスラプション予測器を開発し、それらを高性能領域へ転移させることは、ITERのような大型の炉規模トカマクおよびそれ以降にとって有効な解決策である。本研究では、HL-2Aトカマク実験からの異なるデータサブセットを用いて、多数のモデルを訓練する。モデルに対してSHapley Additive exPlanations(SHAP)解析を実行し、異なる性能領域間には異なる、さらには矛盾するパターンが存在することを明らかにする。したがって、異なる性能領域間でデータを単純に混合しても、最適な結果は得られない。この解析に基づき、我々は最適化された戦略で生成されたデータセットを用いてモデルを訓練する、インスタンスベースの転移学習手法を提案する。この戦略は、SHAPモデル解析によって明らかにされた高・低性能ディスチャージ間の差異の背後にある物理に基づく、インスタンスと特徴の選択を含む。TrAdaBoost手法により、少数の高性能運転データを用いて、モデル性能は0.78 BA(均衡精度)から0.86 BAへと大幅に改善された。