Adebanji Adelowo
← Projects

LLM systems & model compression

Applied ML · Generative AI

My contribution

I developed a collection of LLM projects, including FinBERT pruning and knowledge distillation with saved evaluation results.

Applied ML collection

FinBERT compression: distillation recovers accuracy to within 0.7 points of the baseline after layer dropping; absolute scores are optimistic because the test split likely overlaps the baseline's fine-tuning data.

FinBERT model compression: Teacher model to Compression to Student model
Experiment schematic · financial sentiment classification. This diagram illustrates the documented workflow. Select figure to enlarge.

Problem

Exploring retrieval-augmented generation, structured data access, and model compression through implemented pipelines and evaluations.

System

A collection of nine projects at varying levels of completeness, including a Medical RAG Assistant (LangChain, ChromaDB, conversational memory), a two-stage Enterprise NL-to-SQL pipeline, LoRA/QLoRA parameter-efficient fine-tuning, and a full compression pipeline (Taylor-gradient pruning plus knowledge distillation) applied to a FinBERT financial-sentiment classifier.

Result

On a 453-sentence Financial PhraseBank test split, layer dropping alone cuts the 97.6% baseline accuracy sharply, to 61.8%; knowledge distillation recovers nearly all of it, landing within 0.7 points of baseline while keeping the ~19% reduction in parameters and model size. The baseline (ProsusAI/finbert) was fine-tuned on the Financial PhraseBank, so most test sentences were probably seen in training: the absolute accuracies are optimistic, and the finding is the relative effect of the compression steps.

96.91%Test accuracy, distilled student (test set likely seen by the baseline)
0.9691Test weighted F1 (distilled student)
88.22MParams (vs 110M Baseline)

Scope

The reported metrics apply to the FinBERT compression experiment. Other projects in the collection have different implementation and evaluation coverage; they are not all deployed or fully evaluated systems.