MEGA Hub

Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA

Authors

Do you know Shaswata Mitra?You can claim authorship or link another user.Do you know Subash Neupane?You can claim authorship or link another user.Do you know Trisha Chakraborty?You can claim authorship or link another user.Do you know Himanshu Tripathi?You can claim authorship or link another user.Do you know Sudip Mittal?You can claim authorship or link another user.Do you know Aritran Piplai?You can claim authorship or link another user.Do you know Shahram Rahimi?You can claim authorship or link another user.

Abstract

Large Language Models (LLMs) are increasingly fine-tuned for critical-domain Question-Answering (QA), yet choosing which small model to adapt, before paying the cost of adaptation, remains difficult. Fine-tuning can improve domain alignment, but it may also erode prior knowledge, weaken instruction-following, or increase hallucination, especially when labeled data are scarce or rapidly evolving as in cybersecurity. We present FiT (Find before Fine-Tune), a task-oriented diagnostic framework that characterizes small LLMs along three capabilities required for cybersecurity QA: vocabulary recognition, parametric knowledge, and contextualization of retrieved information. Using FiT, we conduct an empirical study of five open-weight 7-billion-parameter models under two fine-tuning regimes. We find that fine-tuning does not uniformly help: it consistently degrades vocabulary and parametric knowledge in small models, and the two regimes trade off differently. Knowledge-focused tuning causes moderate, rank-preserving degradation, whereas instruction-focused tuning collapses measured knowledge through induced abstention, inverting the knowledge ranking while leaving retrieval-grounded contextualization essentially intact. We quantify these regime-specific patterns with rank-correlation analysis and show that pre-fine-tuning FiT scores anticipate the direction of post-tuning change. Our results suggest that task-oriented diagnosis can screen out unsuitable models, avoid unnecessary fine-tuning, and support safer deployment of small LLMs in cybersecurity QA pipelines.

Community

00

Publication notes

Author note
8 pages, 5 figures, 4 tables, IEEE ICMLA