Publications

Peer-reviewed papers and preprints, newest first.

  1. CovCast: When Will Fuzzing Stop Paying Off? A Specialized Tabular Foundation Model for Coverage-Rate Forecasting

    Amine Lbath

    Preprint, 2026, under review

    Fuzzing finds security vulnerabilities by executing a program on millions of randomly generated and mutated inputs, and a single campaign often runs for days on many machines. As fewer parts of the program remain unexplored, the probability of finding new bugs decreases, and practitioners must decide when to stop a campaign. This decision requires forecasting the coverage rate of a campaign, the number of new code elements it will cover per unit of time in the coming hours or days. Forecasting coverage is challenging because coverage grows irregularly, alternating between long plateaus and sudden increases. Existing forecasting methods adapt statistical estimators whose assumptions fuzzers violate, and analyze each campaign in isolation, which limits their accuracy.

    We present CovCast (Coverage foreCaster), a novel tabular foundation model for forecasting fuzzing coverage. CovCast is pretrained on synthetic fuzzing campaigns generated by a simulator, and forecasts the future coverage of any new campaign from the coverage observed so far, without any re-training. On the FuzzTastic benchmark (133 campaigns on five programs), CovCast reduces the median forecast error of the prior state of the art by 17%, and by 39% for forecasts made on the second and third days of a campaign, when stopping saves the most compute. Stopping campaigns based on its forecasts saves almost four times as much compute as a fixed time budget for the same loss of coverage, as much as the prior state of the art. Unlike existing methods, which require per-element execution counts, CovCast only needs the coverage over time that every fuzzer reports, which makes it applicable to any campaign. On 1,833 FuzzBench campaigns of 11 fuzzers, which do not record the data that existing methods need, CovCast reduces the forecast error of the strongest applicable baseline by 23% and is the most accurate method for every fuzzer and every program.

    @unpublished{lbath2026covcast,  title         = {CovCast: When Will Fuzzing Stop Paying Off? A Specialized Tabular Foundation Model for Coverage-Rate Forecasting},  author        = {Lbath, Amine},  year          = {2026},  note          = {Preprint, under review}}
  2. CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training

    Amine Lbath*, Manan Suri*, Aurelien Delaitre, Vadim Okun, Massih-Reza Amini, Ram D. Sriram, Dinesh Manocha

    *equal contribution

    Submitted July 2026, under review

    Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in real-world software. Generally available agents can already aid attackers, who only need to find one exploitable weakness, while defenders must continuously identify and patch all vulnerabilities across fast-growing codebases. Stronger defensive agents would help close this gap, yet the scarcity of security training data with reproducible build and execution environments remains a bottleneck. We present CyberForge, a framework that synthesizes executable, repository-level security training data by injecting vulnerabilities into real C/C++ projects. It validates each instance dynamically: the injected build must pass the project's unit tests, and generated proof-of-vulnerability (PoV) must trigger on the injected build and not on the clean one. CyberForge is not limited by the availability of disclosed vulnerabilities, therefore it can scale in comparison to data augmentation techniques which rely on historic CVE data. The resulting corpus holds 1034 validated vulnerabilities across 80 projects and 63 weakness categories, with edit locality similar to real CVE patches under a real-versus-real noise floor. Fine-tuning on trajectories collected over this corpus improves SEC-bench patch repair by +3.3 to +14.7 points, in all six configurations of three model scales and two teachers, with the 31B student reaching its GPT-5.4-mini teacher, 72.7% against 74.0%. These gains generalize out of distribution to PatchEval, a corpus containing other programming languages, where every configuration also improves and the 31B student passes its teacher.

    @misc{lbath2026cyberforge,  title         = {CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training},  author        = {Lbath, Amine and Suri, Manan and Delaitre, Aurelien and Okun, Vadim and Amini, Massih-Reza and Sriram, Ram D. and Manocha, Dinesh},  year          = {2026},  eprint        = {2608.06471},  archivePrefix = {arXiv},  primaryClass  = {cs.CR},  url           = {https://arxiv.org/abs/2608.06471}}
  3. A Randomized Framework for Validating Benchmark Claims

    Amine Lbath, Massih-Reza Amini

    Privacy in the Era of Large Opaque Models (PriLOM) Workshop, NeurIPS 2026

    @inproceedings{lbath2026randomized,  title         = {A Randomized Framework for Validating Benchmark Claims},  author        = {Lbath, Amine and Amini, Massih-Reza},  booktitle     = {NeurIPS 2026 Workshop on Privacy in the Era of Large Opaque Models (PriLOM)},  year          = {2026}}
  4. Toward Scalable Automated Repository-Level Datasets for Software Vulnerability Detection

    Amine Lbath

    34th ACM International Conference on the Foundations of Software Engineering (FSE 2026), Doctoral Symposium

    Software vulnerabilities continue to grow in volume and remain difficult to detect in practice. Although learning-based vulnerability detection has progressed, existing benchmarks are largely function-centric and fail to capture realistic, executable, interprocedural settings. Recent repo-level security benchmarks demonstrate the importance of realistic environments, but their manual curation limits scale. This doctoral research proposes an automated benchmark generator that injects realistic vulnerabilities into real-world repositories and synthesizes reproducible proof-of-vulnerability (PoV) exploits, enabling precisely labeled datasets for training and evaluating repo-level vulnerability detection agents. We further investigate an adversarial co-evolution loop between injection and detection agents to improve robustness under realistic constraints.

    @inproceedings{lbath2026toward,  title         = {Toward Scalable Automated Repository-Level Datasets for Software Vulnerability Detection},  author        = {Lbath, Amine},  booktitle     = {Proceedings of the 34th ACM International Conference on the Foundations of Software Engineering (FSE), Doctoral Symposium},  year          = {2026},  publisher     = {ACM},  doi           = {10.1145/3803437.3804870},  url           = {https://dl.acm.org/doi/10.1145/3803437.3804870}}
  5. AVIATOR: Towards AI-Agentic Vulnerability Injection Workflow for High-Fidelity, Large-Scale Code Security Dataset

    Amine Lbath, Massih-Reza Amini, Aurelien Delaitre, Vadim Okun

    Journal of Systems and Software, minor revision

    The increasing complexity of software systems and the sophistication of cyber-attacks have underscored the need for reliable automated software vulnerability detection. Data-driven approaches using deep learning models show promise but critically depend on the availability of large, accurately labeled datasets. Yet existing datasets either suffer from noisy labels, limited vulnerability coverage, or fail to reflect vulnerabilities as they occur in real-world software. This also limits large-scale benchmarking of such solutions. Automated vulnerability injection provides a way to address these limitations, but existing techniques remain limited in coverage, contextual fidelity, or injection success. In this paper, we present AVIATOR, the first AI-agentic vulnerability injection framework. AVIATOR decomposes vulnerability injection into a coordinated workflow of specialized AI agents, tool-based analysis, and iterative self-correction, explicitly mirroring expert reasoning. It integrates RAG and lightweight LoRA-based fine-tuning to produce realistic, category-specific vulnerabilities without relying on handcrafted patterns. Across three benchmarks, AVIATOR achieves high injection fidelity (91-95%) surpassing existing injection techniques in both accuracy and vulnerability coverage. When used for data augmentation to train deep learning-based vulnerability detection (DLVD) models, AVIATOR provides the strongest downstream gains in vulnerability detection. Across models and base datasets, AVIATOR improves average F1 scores by +22% over no augmentation, +25% over VGX, holding the prior best injection success rate, and +3% over VulScribeR, the prior state-of-the-art LLM-based injection model, with +7% higher recall and no precision loss. Its augmented data exhibits the lowest distributional distortion and scales efficiently with <2% syntax rejection at 4.3x lower cost than VulScribeR.

    @misc{lbath2025aviator,  title         = {AVIATOR: Towards AI-Agentic Vulnerability Injection Workflow for High-Fidelity, Large-Scale Code Security Dataset},  author        = {Lbath, Amine and Amini, Massih-Reza and Delaitre, Aurelien and Okun, Vadim},  year          = {2025},  eprint        = {2508.20866},  archivePrefix = {arXiv},  primaryClass  = {cs.CR},  note          = {Minor revision, Journal of Systems and Software},  url           = {https://arxiv.org/abs/2508.20866}}
  6. Energy Efficiency in AI for 5G and Beyond: A DeepRx Case Study

    Amine Lbath, Ibtissam Labriji

    2024 Joint European Conference on Networks and Communications & 6G Summit (EuCNC/6G Summit)

    This study addresses the challenge of balancing energy efficiency with performance in AI/ML models, focusing on DeepRX, a deep learning receiver based on a fully convolutional ResNet architecture. We evaluate the energy consumption of DeepRX, considering factors including FLOPs/Watt and FLOPs/clock, and find consistency between estimated and actual energy usage, influenced by memory access patterns. The research extends to comparing energy dynamics during training and inference phases. A key contribution is the application of knowledge distillation (KD) to train a compact DeepRX student model that emulates the performance of the teacher model but with reduced energy consumption. We experiment with different student model sizes, optimal teacher sizes, and KD hyperparameters. Performance is measured by comparing the Bit Error Rate (BER) performance versus Signal-to-Interference & Noise Ratio (SINR) values of the distilled model and a model trained from scratch. The distilled models demonstrate a lower error floor across SINR levels, highlighting the effectiveness of KD in achieving energy-efficient AI solutions.

    @inproceedings{lbath2024energy,  title         = {Energy Efficiency in AI for 5G and Beyond: A DeepRx Case Study},  author        = {Lbath, Amine and Labriji, Ibtissam},  booktitle     = {2024 Joint European Conference on Networks and Communications \& 6G Summit (EuCNC/6G Summit)},  publisher     = {IEEE},  year          = {2024},  doi           = {10.1109/EuCNC/6GSummit60053.2024.10597065}}