CV

Jie Li

Research Assistant Professor, Department of Computer Science, Texas Tech University

Email: jie.li@ttu.edu · Homepage: lijie.me

Download PDF

RESEARCH INTERESTS

My research focuses on high-performance computing (HPC) and AI infrastructure, with an emphasis on observability, resource management, and energy efficiency. Building on this systems foundation, my research agenda extends to AI agents for HPC operation and management, the security of agents acting within HPC environments, and the integration of HPC with quantum computing.

  • High-performance computing: system monitoring, workload characterization, scheduling, disaggregated memory, and energy-aware resource management.
  • AI infrastructure: LLM serving, KV cache and memory management, and performance and power characterization of AI workloads.
  • AI agents for HPC operations: agent-assisted cluster operation and management, including ongoing development for the NSF REPACSS system.
  • AI agent security in HPC: evaluating and constraining agent behavior in shared computing environments.
  • HPC and quantum computing: quantum arithmetic and polynomial synthesis, with broader interests in hybrid quantum-classical workflows and HPC integration.

EDUCATION

Doctor of Philosophy, Computer Science, Texas Tech University, Lubbock, TX

  • Dissertation: Optimizing High-Performance Computing Systems: Insights from System Monitoring, Workload Management, and Scheduling Strategies. Advisor: Prof. Yong Chen.

Master of Science, Computer Science, Texas Tech University, Lubbock, TX

  • Thesis: PIMS: A Lightweight Processing-in-Memory Accelerator for Stencil Computations. Advisor: Prof. Yong Chen.

Bachelor of Arts, Architecture, Huaqiao University, Xiamen, China

ACADEMIC APPOINTMENTS AND RESEARCH EXPERIENCE

Research Assistant Professor

Department of Computer Science, Texas Tech University, Lubbock, TX
  • Lead research in HPC and AI infrastructure, spanning observability, resource and data management, energy-aware computing, cyberinfrastructure security, and AI-driven system management.
  • Initiated a research direction on AI agents for autonomous HPC operation and management, using REPACSS as a production testbed.
  • Contribute to the operation, optimization, and research use of the NSF REPACSS cluster ($12.25M), including system software integration, performance analysis, and user support.
  • Supervise graduate and undergraduate researchers and develop proposals on secure, autonomous, and energy-efficient computing infrastructure.
  • Co-PI on the NSF REU Site proposal ASPIRE (2026, submitted), which proposes undergraduate research projects in HPC and AI.

Assistant Director

NSF Cloud and Autonomic Computing Center (CAC IUCRC), Texas Tech Site
  • Coordinate faculty, student teams, industry members, and external collaborators in developing and reviewing research projects for this NSF Industry–University Cooperative Research Center.
  • Help organize semiannual Industry Advisory Board meetings, including project planning, presentation preparation, and follow-up with industry participants.
  • Mentor students preparing presentations, posters, and reports for industry review; support proposal development, center reporting, and cross-institutional coordination.

Postdoctoral Researcher

Department of Computer Science, Texas Tech University, Lubbock, TX
  • Co-designed and built the NSF REPACSS cluster ($12.25M) from the ground up, contributing to system architecture, deployment, operation, and performance optimization.
  • Served as technical lead of the Data-Intensive Scalable Computing Laboratory, coordinating research activities, student projects, and collaborative system development.
  • Co-PI on the NSF CICI proposal SHIELD (2025, not funded), which proposed layered defense mechanisms for open-science cyberinfrastructure.

Research Assistant

Data-Intensive Scalable Computing Laboratory, Texas Tech University, Lubbock, TX
  • Across graduate study (2019–2024), published 15 peer-reviewed papers (7 first-author) in venues including CLUSTER, ISC, ICPP, MemSys, IEEE CLOUD, IEEE BigData, and IEEE Transactions on Computers.
  • Developed and maintained MonSTer (later adopted by Dell’s Omnia project), the Disaggregation-Aware Scheduler, and xBGAS simulation tools, supporting research in HPC monitoring, workload characterization, scheduling, and memory systems.
  • Mentored students on HPC monitoring, data management, and workload analysis, including a master’s thesis that led to a CLOUD’23 publication.

Graduate Student Intern

Lawrence Berkeley National Laboratory, Berkeley, CA (Mentors: Brandon Cook and Georgios Michelogiannakis, in John Shalf's group)
  • Designed and implemented pipelines integrating LDMS, DCGM, and Slurm telemetry from NERSC’s Cori and Perlmutter supercomputers.
  • Applied machine learning and deep learning to classify, characterize, and predict HPC job behavior from large-scale time-series telemetry.
  • Led development of the open-source Disaggregation-Aware Scheduler; internship research resulted in first-author publications at ISC’23 and CLUSTER’24.

GRANTS AND PROPOSALS

REPACSS: Empowering Scientific Discovery through Renewable Energy Powered Advanced Computing Systems and Services

National Science Foundation, Category II · Funded · Total award: $12,250,000
  • Role: Contributor to proposal development and infrastructure implementation (PI: Yong Chen; Co-PI: Alan Sill). Contributed proposal sections on data center monitoring and remote control.

ASPIRE: Advancing Scientific Discovery through High-Performance Computing and AI Research

NSF REU Site · Submitted · Requested: $430,000
  • Role: Co-PI (PI: Yong Chen; Co-PI: Maaz Amjad). Designed the undergraduate research projects in HPC and AI and led the proposal’s overall presentation, including its tables and figures.

SHIELD: Strengthening High-Performance Infrastructure with Enhanced Layered Defense

NSF CICI: UCSS · Not funded · Requested: $600,000
  • Role: Co-PI (PI: Yong Chen). Contributed project vision and technical approach for HPC cybersecurity.

DAVinci: An Integrated Data Collection, Automation, and Visualization Framework for HPC Systems

NSF Frameworks · Not funded · Requested: $1,000,000
  • Role: Lead contributor (PI: Yong Chen; Co-PI: Alan Sill). Led core framework design and research methodology.

PEER-REVIEWED PUBLICATIONS

  1. [1]
    X. Wang, A. Tumeo, J. D. Leidel, J. Li, and Y. Chen, “MAC: Memory Access Coalescer for 3D-Stacked Memory,” in Proceedings of the 48th International Conference on Parallel Processing (ICPP’19), 2019, pp. 1–10. doi: 10.1145/3337821.3337867. (Acceptance Rate: 106/405=26.2%)
  2. [2]
    J. Li, X. Wang, A. Tumeo, B. Williams, J. D. Leidel, and Y. Chen, “PIMS: A Lightweight Processing-in-Memory Accelerator for Stencil Computations,” in Proceedings of the International Symposium on Memory Systems (MemSys’19), 2019, pp. 41–52. doi: 10.1145/3357526.3357550.
  3. [3]
    V. Pham, N. Nguyen, J. Li, J. Hass, Y. Chen, and T. Dang, “MTSAD: Multivariate Time Series Abnormality Detection and Visualization,” in 2019 IEEE International Conference on Big Data (BigData’19), IEEE, 2019, pp. 3267–3276. doi: 10.1109/BigData47090.2019.9006559.
  4. [4]
    N. Nguyen, J. Hass, Y. Chen, J. Li, A. Sill, and T. Dang, “RadarViewer: Visualizing the Dynamics of Multivariate Data,” in Practice and Experience in Advanced Research Computing (PEARC’20), 2020, pp. 555–556. doi: 10.1145/3311790.3404538.
  5. [5]
    J. Li, G. Ali, N. Nguyen, J. Hass, A. Sill, T. Dang, and Y. Chen, “MonSTer: An Out-of-the-Box Monitoring Tool for High Performance Computing Systems,” in 2020 IEEE International Conference on Cluster Computing (CLUSTER’20), IEEE, 2020, pp. 119–129. doi: 10.1109/CLUSTER49012.2020.00022. (Acceptance Rate: 27/132=20.5%)
  6. [6]
    X. Wang, A. Tumeo, J. D. Leidel, J. Li, and Y. Chen, “HAM: Hotspot-Aware Manager for Improving Communications With 3D-Stacked Memory,” IEEE Transactions on Computers (IEEE Trans Comput), vol. 70, no. 6, pp. 833–848, 2021, doi: 10.1109/TC.2021.3066982.
  7. [7]
    T. Dang, N. Nguyen, J. Hass, J. Li, Y. Chen, and A. Sill, “The Gap between Visualization Research and Visualization Software in High-Performance Computing Center,” The Gap between Visualization Research and Visualization Software (VisGap’21), 2021, doi: 10.2312/visgap.20211089.
  8. [8]
    T. Dang, N. V. T. Nguyen, J. Li, A. Sill, J. Hass, and Y. Chen, “JobViewer: Graph-based Visualization for Monitoring High-Performance Computing System,” in 2022 IEEE/ACM International Conference on Big Data Computing, Applications and Technologies (BDCAT’22), IEEE, 2022, pp. 110–119. doi: 10.1109/BDCAT56447.2022.00021.
  9. [9]
    J. Li, G. Michelogiannakis, B. Cook, D. Cooray, and Y. Chen, “Analyzing Resource Utilization in an HPC System: A Case Study of NERSC’s Perlmutter,” in International Conference on High Performance Computing (ISC’23), Springer, 2023, pp. 297–316. doi: 10.1007/978-3-031-32041-5_16. (Acceptance Rate: 21/78=26.9%)
  10. [10]
    J. Li, R. Wang, G. Ali, T. Dang, A. Sill, and Y. Chen, “Workload Failure Prediction for Data Centers,” in 2023 IEEE 16th International Conference on Cloud Computing (CLOUD’23), 2023, pp. 479–485. doi: 10.1109/CLOUD60044.2023.00064.
  11. [11]
    C. E. Caon, J. Li, and Y. Chen, “Effective Management of Time Series Data,” in 2023 IEEE 16th International Conference on Cloud Computing (CLOUD’23), 2023, pp. 408–414. doi: 10.1109/CLOUD60044.2023.00055.
  12. [12]
    T. Dang, N. V. T. Nguyen, J. Li, A. Sill, and Y. Chen, “Spiro: Order-Preserving Visualization in High Performance Computing Monitoring,” in International Symposium on Visual Computing (ISVC’23), Springer, 2023, pp. 109–120. doi: 10.1007/978-3-031-47969-4_9.
  13. [13]
    J. Li, J. D. Leidel, B. Page, and Y. Chen, “Towards Cycle-accurate Simulation of xBGAS,” in 2024 International Conference on Computing, Networking and Communications (ICNC’24), IEEE, 2024, pp. 468–472. doi: 10.1109/ICNC59896.2024.10556078.
  14. [14]
    J. Li, G. Michelogiannakis, B. Cook, J. Shalf, and Y. Chen, “Scheduling and Allocation of Disaggregated Memory Resources in HPC Systems,” in 2024 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW’24), IEEE, 2024, pp. 1202–1203. doi: 10.1109/IPDPSW63119.2024.00206.
  15. [15]
    J. Li, G. Michelogiannakis, S. Maloney, B. Cook, E. Suarez, J. Shalf, and Y. Chen, “Job Scheduling in High Performance Computing Systems with Disaggregated Memory Resources,” in 2024 IEEE International Conference on Cluster Computing (CLUSTER’24), IEEE, 2024, pp. 297–309. doi: 10.1109/CLUSTER59578.2024.00033.
  16. [16]
    C. Niu, W. Zhang, J. Li, Y. Zhao, T. Wang, X. Wang, and Y. Chen, “TokenPowerBench: Benchmarking the Power Consumption of LLM Inference,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI’26), 2026, pp. 32582–32590. doi: 10.1609/aaai.v40i38.40535. (Acceptance Rate: 4,167/23,680=17.6%)
  17. [17]
    Y. Zhao, J. Li, C. Niu, A. Sill, and Y. Chen, “Power-Centric Observability for HPC Systems: Design, Deployment, and Evaluation on REPACSS,” in Practice and Experience in Advanced Research Computing (PEARC’26), 2026, pp. 1–8.
  18. [18]
    Y. Yang, M. A. Rafi, J. S. Firoz, Z. Peng, X. Wang, K. Bouasker, S. Ghosh, J. Li, D. Wang, D. Li, H. Jeon, K. Barker, N. Tallent, and L. Guo, “OCEAN: Open-Source CXL Emulation for Hyperscale Architecture and Networking,” in The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC’26), 2026. Accepted

PREPRINTS

  1. [1]
    J. Li, B. Cook, and Y. Chen, “ARcode: HPC Application Recognition Through Image-encoded Monitoring Data,” arXiv preprint arXiv:2301.08612, 2023, doi: 10.48550/arXiv.2301.08612.
  2. [2]
    J. Li, T. Wang, and Y. Chen, “From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving,” arXiv preprint arXiv:2607.02574, 2026.
  3. [3]
    J. Li, “Trusted Credentials, Untrusted Behavior: Benchmarking LLM-Agent Security in High-Performance Computing,” arXiv preprint arXiv:2607.18485, 2026.

MANUSCRIPTS UNDER SUBMISSION OR REVISION

  • Ziqing Guo, Jie Li, Ziwen Pan, and Yong Chen. “NNQA: Neural-Native Quantum Arithmetic for End-to-End Polynomial Synthesis.” Submitted to AAAI-27.
  • Tongyang Wang, Jie Li, Chenxu Niu, and Yong Chen. “ECHO: Online Control of CPU Frequency for Energy-Efficient High-Performance Computing.”
  • Batuhan Sencer, Chenxu Niu, Jie Li, and Yong Chen. “Slimming Models, Saving Watts: Understanding and Modeling the Impact of Knowledge Distillation on GPU Clusters.”

OPEN-SOURCE RESEARCH SOFTWARE

  • MonSTer: Out-of-the-box HPC monitoring framework. Published at CLUSTER’20; adopted by Dell’s Omnia project.
  • Disaggregation-Aware Scheduler: Simulation framework for job scheduling in HPC systems with disaggregated memory. Used in CLUSTER’24 and IPDPSW’24 publications.
  • xBGAS REV-CPU: Cycle-accurate xBGAS simulation extension based on REV-CPU. Developed with Tactical Computing Laboratories; published at ICNC’24.
  • xBGAS Runtime: Lightweight runtime supporting global address space extensions for RISC-V-based xBGAS simulation. Developed with Tactical Computing Laboratories.

TEACHING

Teaching interests: operating systems, parallel and high-performance computing, computer architecture, distributed and cloud computing, and systems for machine learning at the undergraduate and graduate levels; new graduate seminars on AI infrastructure and on AI agents for computing systems.

Invited Lecturer and Programming Project Designer

Parallel Processing (graduate), Texas Tech University
  • Delivered four invited lectures and designed programming projects for 27 students; topics included job scheduling, compilation and job submission, and OpenMP.

RESEARCH MENTORING

Graduate Students

  • Jaechang Kim, M.S. student. AI agents for REPACSS cluster operation and management, automating Warewulf, Ansible, Spack, and Slurm workflows.
  • Rupak Kadel, Ph.D. student. Improving HPC monitoring for anomaly and intrusion detection.
  • Cristiano Caon, M.S. Data volume reduction and query optimization in time-series databases. Outcome: CLOUD’23 publication. Independent Study (CS 7000) and Master's Thesis
  • Aniruddh Sanjaysinh Chavda (M.S.) and Huyen Nguyen (Ph.D.). Usage behavior analysis by clustering job accounting data. Advanced Operating Systems
  • Ruonan Wu, M.S. Job accounting data analysis for the Quanah cluster. Advanced Operating Systems
  • Ashhrita Puradamane Balachandra, M.S. Improving InfluxDB query performance. Advanced Operating Systems

Undergraduate Students

  • Yusheng Han and Zachary Kay. Running HPC applications and analyzing performance on the RedRaider cluster. Independent Study (CS 4000)
  • Casey Root. Monitoring queue status through the Slurm REST API. Independent Study (CS 4000)

PRESENTATIONS

  • Integrated Data Collection and Visualization Framework for Data Centers based on Telemetry Model. NSF CAC Industry Advisory Board Conferences (Lubbock, TX; Denton, TX; Tucson, AZ; Lubbock, TX)
  • Towards Cycle-Accurate Simulation of xBGAS. Latch-Up 2024, Cambridge, MA
  • Workload Failure Prediction for Data Centers. IEEE CLOUD'23, Chicago, IL
  • Advanced Visualization and Data Analysis of HPC Cluster and User Application Behavior. ACM/IEEE SC'21
  • MonSTer: An Out-of-the-Box Monitoring Tool for High Performance Computing Systems. IEEE CLUSTER'20
  • PIMS: A Lightweight Processing-in-Memory Accelerator for Stencil Computations. MemSys'19

REFEREED POSTERS

  • J. Li, B. Cook, G. Michelogiannakis, and Y. Chen. A Holistic View of Memory Utilization on Perlmutter. SC'22
  • J. Li, B. Cook, and Y. Chen. Detecting and Identifying Applications by Job Signatures. SC'21
  • X. Wang, J. Li, A. Tumeo, J. D. Leidel, and Y. Chen. Memory Hotspot Optimizations for 3D-Stacked Memory. PACT'19

AWARDS AND HONORS

  • Best Poster Award, NSF Cloud and Autonomic Computing Center Industry Advisory Board Conference
  • Summer Thesis/Dissertation Research Award ($2,300), Texas Tech University Graduate School

PROFESSIONAL SERVICE

  • Program committee: NeurIPS 2026; AAAI 2026.
  • Conference reviewer: IEEE ISCAS 2026; IEEE BigData 2022, 2023, 2025; IEEE/ACM CCGrid 2024.
  • Journal reviewer: IEEE Computer Architecture Letters (2025); The Journal of Supercomputing (2023).
  • Conference sub-reviewer: IEEE IPDPS 2023; IEEE ICDCS 2022; ACM/IEEE SC 2022; International Parallel Data Systems Workshop (PDSW) 2022; IEEE Smart Data Services 2020.
  • Student volunteer: ACM/IEEE SC’21, St. Louis, MO; ACM/IEEE SC’19, Denver, CO.

Last updated: September 2026