Recently, with the advancement of the AI field, reinforcement learning (RL) has increasingly been applied to plasma control on tokamak devices. However, possibly due to the generally high training costs of reinforcement learning based on first-principle physical models and the uncertainty in ensuring simulation results align perfectly with tokamak experiments, feedback control experiments using reinforcement learning specifically for plasma kinetic parameters on tokamaks remain scarce. To address this challenge, this work proposes a novel design scheme including the development of a low computational cost environment. This environment is derived from EAST modulation experiments data through system identification. To tackle issues of noise and actuator limitations encountered in experiments, data preprocessing methods were employed. During training, the agent collected data across multiple plasma scenarios to update its strategy, and the performance of the RL controller was fine-tuned by adjusting the weight of the integral term of the error in the reward function. The effectiveness and robustness of the proposed design were then validated in a simulated environment. Finally, the scheme was successfully implemented on EAST, effectively tracking the βp target with lower hybrid wave (LHW) at 4.6 GHz as the actuator, and providing reference for implementing feedback control based on reinforcement learning in tokamaks.
近年、AI分野の進展に伴い、強化学習(RL)はトカマク装置におけるプラズマ制御にますます応用されている。しかしながら、第一原理物理モデルに基づく強化学習の一般的に高い訓練コストと、シミュレーション結果がトカマク実験と完全に一致することを保証することの不確実性のためか、トカマクにおけるプラズマ運動論的パラメータに特化した強化学習を用いたフィードバック制御実験は依然として稀である。この課題に対処するため、本研究では、低計算コスト環境の開発を含む新しい設計手法を提案する。この環境は、システム同定を通じてEAST変調実験データから導出される。実験で遭遇するノイズとアクチュエータの制限の問題に取り組むため、データ前処理手法が用いられた。訓練中、エージェントは複数のプラズマシナリオにわたってデータを収集して方策を更新し、RLコントローラの性能は報酬関数における誤差の積分項の重みを調整することで微調整された。提案された設計の有効性と堅牢性は、その後シミュレーション環境で検証された。最後に、この手法はEASTで正常に実装され、4.6 GHzの低ハイブリッド波(LHW)をアクチュエータとしてβp目標値を効果的に追跡し、トカマクにおける強化学習に基づくフィードバック制御の実装の参考を提供する。