LAET: A Layer-Wise Adaptive Ensemble Tuning Framework for Pretrained Language Models
Jawad Ibn Ahad, Muhammad Rafsan Kabir, Robin Krambroeckers, Sifat Momen, Nabeel Mohammed, Shafin Rahman
β Accepted at IEEE International Conference on Big Data 2025
π Paper: arXiv:2511.11315 Β· IEEE Xplore
Abstract
We propose LAET (Layer-wise Adaptive Ensemble Tuning), a novel fine-tuning strategy for pretrained LLMs that identifies and fine-tunes only the most task-relevant layers while freezing less critical ones. By probing hidden state representations per layer through a lightweight classifier, LAET reduces trainable parameters by up to 60% while improving task-specific performance. Ensemble predictions across selected layers further improve robustness. Across 23 financial NLP benchmarks spanning textual analysis, risk management, and forecasting, LAET (on ~3B models) consistently outperforms GPT-4, GPT-4o, DoRA, LoRA, and specialized financial LLMs.
Introduction
Fine-tuning all parameters of large pretrained LLMs is computationally prohibitive and often unnecessary β evidence shows that only a fraction of layers contribute meaningfully to any downstream task. Existing PEFT methods (LoRA, DoRA, AdaLoRA) apply uniform or heuristically distributed parameter budgets, ignoring per-layer utility. LAET addresses this by probing each layer independently to score task relevance, then selectively fine-tuning only high-value layers.
Models used (β³β, β³β, β³β):
- β³β: Gemma-2-2B (2B parameters)
- β³β: Llama-3.2-3B (3.2B parameters)
- β³β: Phi-3.5-mini (3.8B parameters)
π― Key Contributions
- Layer-wise Probing β A shared lightweight classifier trained on frozen per-layer hidden states scores each layerβs task relevance via accuracy (mβ) and F1 (mβ).
- Adaptive Layer Selection β Dynamic thresholds
Ξ΄ = Ξ±Β·Οbased on score standard deviations automatically select the best-performing layers. - Ensemble Voting β Majority vote over predictions from all selected layers reduces sensitivity to individual layer noise.
- 60% Parameter Savings β Frozen layers do not receive gradient updates; only selected layers are fine-tuned.
- Outperforms GPT-4 β ~3B parameter models with LAET beat GPT-4 and GPT-4o on financial NLP benchmarks.
Introduction

Figure 1: Comparison of full fine-tuning, parameter-efficient fine-tuning (PEFT), and LAET. Full fine-tuning updates all layers; PEFT applies uniform low-rank adapters. LAET identifies and fine-tunes only the most task-relevant layers, leaving the rest frozen β reducing trainable parameters by up to 60% while improving task performance.
Methodology

Figure 2: LAET methodology. (a) For each layer l in pretrained model β³, hidden state representations r_i^(l) are extracted from frozen layers; a shared lightweight classifier β±_Ο computes logits z_i^l and cross-entropy loss β_l independently per layer. (b) The best-performing layers are selected based on evaluation metrics (accuracy mβ, F1 mβ) and their deviations from maximum values using adaptive threshold Ξ΄ = Ξ±Β·Ο. Only selected layers β¬ are unfrozen for fine-tuning; inference uses majority voting across all selected layers.
Problem Formulation
Let β³ be a pretrained LLM with L layers. For input tokens x, each layer l produces hidden states H_l = {hββ½Λ‘βΎ, β¦, hββ½Λ‘βΎ}.
The last-token representation per layer: r_l = h_n^(l)[:, r] β βα΅
A lightweight classifier β±_Ο with shared weights is trained independently on each layerβs frozen representations.
Layer Probing

Figure 3: Layer-wise probing evaluation across three hidden state representation strategies: Last Token (LT), Sum of All Tokens (SaT), and Average of All Tokens (AvT). Results show which strategy produces the most task-discriminative representations per layer for adaptive selection.
Per-layer cross-entropy loss:
ββ = -(1/N) Ξ£α΅’ Ξ£β±Ό π«[yα΅’=j] Β· log(pα΅’Λ‘[j])
where pα΅’Λ‘ = softmax(β±_Ο(rα΅’β½Λ‘βΎ)) β βα΅.
Adaptive Layer Selection
For performance metrics mβ (accuracy), mβ (F1) per layer:
Οββ = std({mβΒΉ,...,mβα΄Έ}), Οββ = std({mβΒΉ,...,mβα΄Έ})
Ξ΄ββ = Ξ±Β·Οββ, Ξ΄ββ = Ξ²Β·Οββ
Layer l is selected into set β¬ if no other layer lβ satisfies: mβΛ‘' β₯ mβΛ‘ + Ξ΄ββ AND mβΛ‘' β₯ mβΛ‘ + Ξ΄ββ
Fine-tuning and Ensemble
Only layers l β β¬ are unfrozen. Training objective:
β_β¬ = (1/|β¬|) Ξ£βββ¬ ββ(ΞΈβ, Ο)
Inference β majority voting:
Ε· = argmax_{cβπ} Ξ£βββ¬ π(Ε·β = c)
Ensemble error bound (conditional independence, average error Ξ΅Μ):
π«(Ε· β y) β€ exp(β2|β¬|(0.5 β Ξ΅Μ)Β²)

Figure 4: Ablation study comparing LAET layer selection methods across four datasets. Results validate that the adaptive threshold selection (Ξ΄ = Ξ±Β·Ο) outperforms static or random layer selection, confirming that task-relevant layer identification is the key driver of LAETβs performance gains.
π Table I β Dataset Statistics (23 Financial NLP Benchmarks)
| Dataset | Task | Language | Type | Train | Test | Metric |
|---|---|---|---|---|---|---|
| FPB | Sentiment Analysis | EN | news | 3,100 | 970 | Acc, F1 |
| FiQA-SA | Sentiment Analysis | EN | news/tweets | 750 | 235 | Acc, F1 |
| TSA | Sentiment Analysis | EN | headlines | 448 | 113 | RMSE |
| TSA | Sentiment Analysis | ES | headlines | 3,063 | 766 | Acc, F1 |
| FinanceES | Sentiment Analysis | ES | headlines | 5,084 | 1,272 | Acc, F1 |
| Headlines | News Classification | EN | headlines | 16,437 | 4,110 | Acc, Avg-F1 |
| FOMC | Hawkish-Dovish | EN | news | 396 | 100 | Acc, F1 |
| FinArg-ACC | Argument Classification | EN | comments | 6,202 | 1,551 | Acc, MiF1 |
| MultiFin-E | Multi-class | EN | headlines | 5,350 | 1,340 | Acc, MiF1 |
| MultiFin-S | Multi-class | ES | headlines | 210 | 20 | Acc, MiF1 |
| MA | Deal Classification | EN | news/tweets | 400 | 100 | Acc, MiF1 |
| BigData22 | Stock Movement | EN | tweets+prices | 4,900 | 1,470 | Acc, F1, MCC |
| ACL18 | Stock Movement | EN | tweets+prices | 20,800 | 3,720 | Acc, F1, MCC |
| CIKM18 | Stock Movement | EN | tweets+prices | 3,400 | 431 | Acc, F1, MCC |
| German | Credit Scoring | EN | biography | 700 | 200 | Acc, F1, MCC |
| Australian | Credit Scoring | EN | biography | 482 | 139 | Acc, F1, MCC |
| LendingClub | Credit Scoring | EN | loan record | 9,420 | 2,690 | Acc, F1, MCC |
| ccf | Fraud Detection | EN | financial record | 7,970 | 2,280 | Acc, F1, MCC |
| cfraud | Fraud Detection | EN | financial record | 7,340 | 2,100 | Acc, F1, MCC |
| polish | Financial Distress | EN | financial profile | 6,080 | 1,740 | Acc, F1, MCC |
| taiwan | Financial Distress | EN | financial profile | 4,770 | 1,370 | Acc, F1, MCC |
| ProtoSegno | Financial Distress | EN | financial profile | 8,330 | 2,380 | Acc, F1, MCC |
| travrealtins. | Claim Analysis | EN | financial profile | 8,870 | 2,530 | Acc, F1, MCC |
π Table II β Textual Analysis (TA) Results
| Category | Model | FPB (Acc/F1) | FiQA (Acc/F1) | TSA-E (Acc/F1) | TSA-S (Acc/F1) | Fin.ES (Acc) | Head. (Acc/F1) | FOMC (Acc) | FinArg (Acc/MiF1) | MultiFin-E | MultiFin-S | MA (Acc) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| General LLM | ChatGPT | 0.78/0.78 | β/0.60 | 0.21/0.24 | 0.13/0.08 | β | 0.48/0.47 | β | β | β | β | β |
| Β | GPT-4 | 0.76/0.78 | β/0.80 | 0.47/0.56 | 0.15/0.09 | β | 0.60/0.60 | β | β | β | β | β |
| Β | Gemini | 0.77/0.77 | β/0.81 | β | β | β | 0.62/0.62 | β | 0.31/β | β/β | 0.84/β | β |
| Β | LLaMA2-7B | 0.68/0.65 | β/0.77 | 0.07/0.04 | 0.14/0.13 | β | 0.23/0.11 | β | 0.46/β | β | 0.70/β | β |
| Β | LLaMA2-70B | 0.73/0.72 | β/0.83 | β | β | β | 0.63/0.63 | β | 0.58/β | β | 0.86/β | β |
| Β | LLaMA3-8B | 0.52/0.52 | β/0.70 | β | β | β | 0.39/0.39 | β | 0.51/β | β | 0.34/β | β |
| Few-shot | GPT-4o | 0.87/0.88 | 0.70/0.87 | 0.72/0.66 | 0.70/0.58 | 0.65 | 0.72/0.65 | 0.70 | 0.61/β | 0.56/β | β | β |
| Β | Deepseek V3 | 0.83/0.83 | 0.70/0.85 | 0.81/0.81 | 0.11/0.11 | 0.65 | 0.66/0.99 | 0.99 | 0.72/β | 0.65/β | β | β |
| Β | Qwen2.5 72B | 0.73/0.73 | 0.59/0.83 | 0.76/0.78 | 0.13/0.13 | 0.67 | 0.60/0.91 | 0.91 | 0.60/β | 0.55/β | β | β |
| Financial LLM | FinMA-7B | 0.88/0.88 | β/0.84 | β | β | β | 0.98/β | β | β | β | β | β |
| Β | FinMA-30B | 0.87/0.88 | β/0.87 | β | β | β | 0.97/β | β | β | β | β | β |
| Β | FinMA-ESB | 0.83/0.83 | β/0.85 | 0.85/0.86 | 0.11/0.11 | β | 0.96/0.55 | 0.49 | β | 0.99/0.99 | β | β |
| PEFT | β³β-LoRA | 0.84/0.83 | 0.84/0.83 | 0.80/0.79 | 0.75/0.74 | 0.92 | 0.92/0.66 | 0.65 | 0.71/0.70 | 0.86/0.86 | 0.83/0.82 | 0.82 |
| Β | β³β-LoRA | 0.85/0.85 | 0.84/0.84 | 0.80/0.79 | 0.70/0.70 | 0.92 | 0.92/0.66 | 0.66 | 0.72/0.72 | 0.86/0.86 | 0.88/0.84 | 0.84 |
| Β | β³β-LoRA | 0.80/0.81 | 0.86/0.86 | 0.80/0.79 | 0.72/0.71 | 0.93 | 0.93/0.66 | 0.66 | 0.68/0.68 | 0.84/0.83 | 0.89/0.85 | 0.85 |
| Β | β³β-DoRA | 0.85/0.84 | 0.85/0.84 | 0.81/0.81 | 0.77/0.76 | 0.94 | 0.94/0.66 | 0.66 | 0.73/0.72 | 0.87/0.87 | 0.84/0.83 | 0.83 |
| Β | β³β-DoRA | 0.86/0.86 | 0.85/0.85 | 0.81/0.81 | 0.72/0.72 | 0.94 | 0.94/0.68 | 0.68 | 0.74/0.74 | 0.87/0.87 | 0.90/0.85 | 0.85 |
| Β | β³β-DoRA | 0.86/0.86 | 0.86/0.86 | 0.81/0.81 | 0.74/0.73 | 0.94 | 0.94/0.68 | 0.68 | 0.70/0.70 | 0.85/0.84 | 0.91/0.87 | 0.87 |
| Β | β³β-AdaLoRA | 0.81/0.80 | 0.81/0.80 | 0.77/0.76 | 0.70/0.68 | 0.90 | 0.90/0.63 | 0.62 | 0.68/0.66 | 0.82/0.82 | 0.79/0.78 | 0.78 |
| Β | β³β-AdaLoRA | 0.82/0.82 | 0.82/0.82 | 0.77/0.76 | 0.67/0.67 | 0.91 | 0.91/0.64 | 0.63 | 0.70/0.70 | 0.83/0.82 | 0.84/0.81 | 0.81 |
| Β | β³β-AdaLoRA | 0.81/0.82 | 0.83/0.83 | 0.77/0.76 | 0.69/0.69 | 0.91 | 0.91/0.65 | 0.64 | 0.66/0.66 | 0.81/0.80 | 0.85/0.82 | 0.82 |
| LAET (Ours) | β³β-LAET | 0.88/0.87 | 0.88/0.87 | 0.84/0.83 | 0.79/0.78 | 0.97 | 0.97/0.70 | 0.68 | 0.75/0.74 | 0.90/0.90 | 0.87/0.86 | 0.86 |
| Β | β³β-LAET | 0.89/0.89 | 0.88/0.88 | 0.84/0.83 | 0.74/0.74 | 0.97 | 0.97/0.70 | 0.70 | 0.76/0.76 | 0.90/0.90 | 0.93/0.88 | 0.88 |
| Β | β³β-LAET | 0.89/0.89 | 0.90/0.90 | 0.84/0.83 | 0.76/0.75 | 0.98 | 0.98/0.70 | 0.69 | 0.72/0.72 | 0.88/0.87 | 0.94/0.89 | 0.89 |
π Table III β Risk Management (RM) Results
| Category | Model | German (Acc/F1/MCC) | Australian (Acc/F1/MCC) | LendingClub (Acc/F1/MCC) | ccf (Acc/F1/MCC) | cfraud (Acc/F1/MCC) | polish (Acc/F1/MCC) | taiwan (Acc/F1/MCC) |
|---|---|---|---|---|---|---|---|---|
| General LLM | GPT-4 | 0.55/0.51/β | 0.74/0.75/β | β/β/β | β/β/β | β/β/β | β/β/β | β/β/β |
| Few-shot | GPT-4o | 0.61/0.73/0.53 | 0.86/0.84/0.51 | 0.62/0.51/0.65 | 0.59/0.65/0.73 | 0.51/0.58/0.61 | 0.58/0.51/0.67 | 0.73/0.66/0.62 |
| Β | Deepseek V3 | 0.68/0.57/0.58 | 0.59/0.59/0.53 | 0.55/0.65/0.53 | 0.50/0.65/0.56 | 0.68/0.62/0.68 | 0.47/0.51/0.71 | 0.50/0.48/0.71 |
| Β | Qwen2.5 72B | 0.60/0.65/0.66 | 0.59/0.57/0.57 | 0.65/0.58/0.45 | 0.45/0.50/0.51 | 0.62/0.59/0.66 | 0.46/0.48/0.65 | 0.54/0.71/0.51 |
| PEFT | β³β-DoRA | 0.68/0.56/0.00 | 0.81/0.82/0.68 | 0.93/0.93/0.87 | 0.95/0.93/0.00 | 0.89/0.88/0.00 | 0.90/0.87/0.00 | 0.91/0.89/0.00 |
| Β | β³β-DoRA | 0.70/0.55/0.00 | 0.82/0.82/0.69 | 0.91/0.91/0.86 | 0.96/0.95/0.00 | 0.90/0.90/0.00 | 0.91/0.88/0.02 | 0.93/0.91/0.02 |
| Β | β³β-DoRA | 0.68/0.55/0.02 | 0.80/0.80/0.68 | 0.92/0.92/0.87 | 0.94/0.94/0.00 | 0.88/0.91/0.00 | 0.90/0.88/0.00 | 0.93/0.93/0.00 |
| Β | β³β-AdaLoRA | 0.70/0.56/0.00 | 0.83/0.79/0.68 | 0.94/0.92/0.87 | 0.94/0.94/0.02 | 0.92/0.90/0.00 | 0.92/0.86/0.00 | 0.93/0.91/0.00 |
| Β | β³β-AdaLoRA | 0.71/0.54/0.00 | 0.80/0.82/0.66 | 0.91/0.94/0.86 | 0.95/0.93/0.00 | 0.89/0.92/0.00 | 0.89/0.89/0.01 | 0.92/0.92/0.01 |
| LAET (Ours) | β³β-LAET | 0.74/0.61/0.00 | 0.88/0.88/0.77 | 0.98/0.98/0.93 | 1.00/1.00/0.00 | 0.96/0.95/0.50 | 0.97/0.95/0.00 | 0.97/0.96/0.00 |
| Β | β³β-LAET | 0.74/0.62/0.00 | 0.87/0.86/0.73 | 0.98/0.98/0.93 | 1.00/1.00/0.00 | 0.95/0.95/0.02 | 0.97/0.94/0.00 | 0.97/0.96/0.00 |
| Β | β³β-LAET | 0.74/0.63/0.00 | 0.87/0.87/0.74 | 0.98/0.98/0.94 | 1.00/1.00/0.00 | 0.96/0.95/0.31 | 0.95/0.92/0.00 | 0.97/0.96/0.00 |
π Table IV β Forecasting (FO) Results
| Category | Model | BigData22 (Acc/F1/MCC) | ACL18 (Acc/F1/MCC) | CIKM18 (Acc/F1/MCC) |
|---|---|---|---|---|
| General LLM | ChatGPT | 0.53/β/β0.025 | 0.50/β/0.005 | 0.55/β/0.005 |
| Β | GPT-4 | 0.54/β/0.030 | 0.52/β/0.020 | 0.57/β/0.020 |
| Β | Gemini | 0.55/β/0.040 | 0.52/β/0.040 | 0.54/β/0.020 |
| Β | LLaMA2-7B | 0.51/β/0.030 | 0.51/β/0.010 | 0.47/β/β0.070 |
| Β | LLaMA3-8B | 0.55/β/0.020 | 0.52/β/0.020 | 0.57/β/0.030 |
| Few-shot | GPT-4o | 0.53/0.54/0.46 | 0.48/0.48/0.019 | 0.45/0.49/0.50 |
| Β | Deepseek V3 | 0.51/0.50/0.49 | 0.45/0.51/0.017 | 0.50/0.51/0.49 |
| Β | Qwen2.5 72B | 0.53/0.53/0.47 | 0.50/0.45/0.025 | 0.51/0.52/0.48 |
| Financial LLM | FinMA-7B | 0.51/β/0.020 | 0.51/β/0.030 | 0.50/β/0.080 |
| Β | FinMA-30B | 0.47/β/0.040 | 0.49/β/0.000 | 0.43/β/β0.050 |
| PEFT | β³β-LoRA | 0.480/0.490/0.000 | 0.430/0.380/0.000 | 0.490/0.380/0.000 |
| Β | β³β-LoRA | 0.510/0.530/0.010 | 0.460/0.350/0.020 | 0.496/0.372/0.000 |
| Β | β³β-LoRA | 0.480/0.460/0.007 | 0.477/0.388/0.003 | 0.501/0.362/0.019 |
| Β | β³β-DoRA | 0.510/0.357/0.000 | 0.466/0.351/0.015 | 0.510/0.380/0.000 |
| Β | β³β-DoRA | 0.482/0.355/0.000 | 0.458/0.380/0.000 | 0.488/0.390/0.000 |
| Β | β³β-DoRA | 0.496/0.334/0.000 | 0.479/0.375/0.000 | 0.519/0.380/0.000 |
| Β | β³β-AdaLoRA | 0.508/0.360/0.015 | 0.483/0.361/0.001 | 0.487/0.375/0.000 |
| Β | β³β-AdaLoRA | 0.494/0.354/0.000 | 0.466/0.369/0.014 | 0.485/0.395/0.000 |
| Β | β³β-AdaLoRA | 0.486/0.350/0.000 | 0.487/0.359/0.000 | 0.500/0.374/0.016 |
| LAET (Ours) | β³β-LAET | 0.550/0.400/0.000 | 0.520/0.420/0.017 | 0.580/0.430/0.000 |
| Β | β³β-LAET | 0.560/0.460/0.060 | 0.550/0.520/0.097 | 0.550/0.550/0.105 |
| Β | β³β-LAET | 0.570/0.550/0.100 | 0.530/0.510/0.000 | 0.580/0.560/0.124 |
π Citation
@inproceedings{ahad2025laet,
title={LAET: A Layer-Wise Adaptive Ensemble Tuning Framework for Pretrained Language Models},
author={Ahad, Jawad Ibn and Kabir, Muhammad Rafsan and Krambroeckers, Robin and Momen, Sifat and Mohammed, Nabeel and Rahman, Shafin},
booktitle={IEEE International Conference on Big Data (BigData)},
year={2025}
}
