Offline reinforcement learning-driven feedforward control in sequencing batch reactor for TMAH-rich semiconductor wastewater

  • Kim, SangYoun
  • Woo, TaeYong
  • Jeong, ChanHyeok
  • Jeon, NaYoung
  • Jun, Byung-Moon
  • ... Heo, SungKu
  • 외 2명
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Semiconductor manufacturing discharges wastewater rich with organic and inorganic contaminants, particularly tetramethylammonium hydroxide (TMAH). TMAH is highly toxic and inhibits nitrification, thereby conventional biological wastewater treatment is interrupted by TMAH and requires advanced strategies to maintain the operational performance. This study proposes an offline reinforcement learning (RL)-based feedforward operation system, enabling data-driven decision-making from historical operation data with minimal human intervention. Firstly, four influent scenarios (High TMAH, Low TMAH, High N, and High C) were identified through fuzzy C-means clustering. Then, the semiconductor activated sludge model (e-ASM) was modified to incorporate the inhibitory effects of TMAH on nitrification process. QMIX algorithm was employed to collect offline datasets, and implicit diffusion Q-learning (IDQL) of offline RL algorithm was employed to train policies under diverse influent conditions. The results showed that IDQL-based feedforward operation systems improved the operating cost index, effluent quality index, and time-weighted average of TMAH by 1.3%, 3.2%, and 9.1%, respectively, compared with fixed-schedule operations under the identified influent scenarios. Converse to fixed schedule operation that relies on fixed aeration and EC dosing, IDQL-based feedforward operation adaptively adjusted operating conditions for each influent scenario. Specifically, RL agents increased the anaerobic duration (18 h) for the High TMAH scenario to mitigate nitrification inhibition, while in the High N scenario, it increased the DO set-point (5.5 mg/L) to ensure complete ammonia oxidation. These findings highlight that offline RL can be utilized as a robust decision-support tool for high-variability wastewater treatment and the foundation for deployment in full-scale industrial wastewater treatment plants.

키워드

Autonomous operationDiffusion modelOffline reinforcement learningSemiconductor wastewaterSequencing batch reactor (SBR)TETRAMETHYLAMMONIUM HYDROXIDE TMAHOPTIMIZATIONTOXICITYMODEL
제목
Offline reinforcement learning-driven feedforward control in sequencing batch reactor for TMAH-rich semiconductor wastewater
저자
Kim, SangYounWoo, TaeYongJeong, ChanHyeokJeon, NaYoungJun, Byung-MoonNam, KiJeonHeo, SungKuYoo, ChangKyoo
DOI
10.1016/j.jwpe.2026.109875
발행일
2026-05
유형
Article
저널명
Journal of Water Process Engineering
87