The primary goal of this project is to provide a robust framework for extracting and analyzing code snippets from the Linux Kernel. These snippets are used for training CodeBERT, a pre-trained model designed for programming languages, to better understand and work with Linux Kernel code.