ABDUL ALI MAMNUN · OPEN-SOURCE CONTRIBUTOR

Open-source contributions.
Merged upstream.

Bug fixes and regression tests contributed to sktime, Keras, and MLflow. Four merged pull requests that strengthen forecasting integrations, tensor operations, and model tracing.

GitHub repository stars · checked September 25, 2026

MY PROJECTS · THREE AREAS OF WORK

From experiments
to working systems.

Explore the areas below, then read the full implementation and results for each project.

01

Scientific & adaptive ML

Predicting physical systems, calibrating quantum controls, and adapting language models. These projects connect model behavior to measurable experiments and inspectable results.

Explore these projects ↓
02

ML infrastructure

The software around a model: validating a release, deploying and monitoring it, recovering from failure, and tracking model dependencies through their lifecycle.

Explore these projects ↓
03

Time-series & forecasting

Learning from data that changes over time. These projects investigate momentum strategies and electricity-demand forecasts, with attention to causal timing, validation, and adaptation.

Explore these projects ↓
THE FULL PROJECT STORIES

What I built.
How I tested it.
What I learned.

Scientific & adaptive ML
THE FULL PROJECT · Scientific ML

FlowBench

A compact scientific machine-learning system that predicts a paired future vorticity field from a two-dimensional fluid snapshot. FlowBench connects data inspection, model training, held-out evaluation, inference serving, and visual inspection in one reproducible pipeline.

32 × 32 VORTICITY FIELDSaved held-out sample
Height & color = field value
Recorded fields, not live inference. Reference and prediction share a scale; error uses its own symmetric scale. Data ↗
01

The problem

A field-level prediction can look plausible while getting important regions or physical summary statistics wrong. FlowBench therefore compares the numerical reference, a neural prediction, and their error, and evaluates the model against persistence—the simple baseline that predicts no change. Vorticity describes local rotation in the fluid; in the visual, height and the blue-to-orange scale both encode the signed field value.

02

Data and model

The project uses the NeuralOperator Team’s Navier–Stokes dataset from Zenodo, working with 32 × 32 fields obtained by stride-four decimation of the 128 × 128 archive. A four-layer residual CNN with 32 channels and 19,105 parameters is trained for 50 epochs; validation MSE selects the checkpoint. Circular padding follows an observed wrap-around continuity check. Splits, normalization, configuration, and data hashes travel with the saved checkpoint.

03

Measured results

On 2,000 held-out pairs, mean relative L2 error is 0.4667 for the CNN versus 0.8359 for persistence. Mean absolute error falls from 0.4711 to 0.2570; mean relative enstrophy error falls from 0.9527 to 0.2272. A separate high-vorticity slice contains 204 test samples, with CNN relative L2 0.4681. The recorded batch-one inference path has p50 0.58 ms and p95 0.72 ms on the reported Apple M5/MPS setup, excluding HTTP overhead.

04

From experiment to service

A single YAML configuration drives inspect → prepare → train → evaluate → serve. The FastAPI endpoint validates the input field and returns the prediction with model version, device, shape, and latency. The Streamlit viewer supports field and error inspection. Checkpoint reload tests verify that evaluation metrics reproduce within tolerance. The CNN is also packaged into ModelLaunch, connecting the scientific experiment to a gated local deployment workflow.

05

What the evidence does not establish

The dataset does not document the prediction horizon, physical units, or domain size sufficiently to claim them here. Enstrophy is computed as a grid mean, not a domain integral. No numerical solver is timed, so these results do not establish solver speed-up. The benchmark is one training run, at one working resolution, with decimation rather than anti-alias filtering. The website replays saved held-out fields; it does not run a new fluid simulation.

Back to project categories ↑
Scientific & adaptive ML
THE FULL PROJECT · Quantum-control simulation

PulseTune

A quantum-control calibration experiment: find a two-segment pulse that implements an Rx(π/2) gate despite detuning and amplitude error, while spending at most 40 objective calls. The system compares Bayesian optimization with simpler alternatives under a documented held-out protocol.

How often did calibration reach the target?

Gate infidelity ≤ 0.001 · 27 held-out runs per method

GP-EI · random start25 / 27
GP-EI · nominal start*27 / 27
Powell9 / 27
Uncalibrated pulse3 / 27
Random search0 / 27

GP-EI handles detuning more reliably. Powell excels when detuning is zero.

40-call ceiling, warm starts included. Nominal is evaluated once. *Nominal-start GP-EI was added after the first held-out run.

View final error and calls-to-target
MethodMedian final infidelity ↓Median calls to target
GP-EI · random start2.52e-430
GP-EI · nominal start*2.41e-428
Powell3.45e-35
Uncalibrated pulse1.21e-21
Random search2.59e-2Not reached
Simulated single-qubit control. The linked demo replays recorded runs. Source ↗
Open the full calibration replay ↗
01

Why calibration is a search problem

A nominal pulse assumes an ideal qubit. Detuning changes the effective frequency, while amplitude error changes control strength. PulseTune simulates those errors and measures average gate infidelity—the mismatch between the resulting operation and the target gate. Lower is better. In a hardware setting, each candidate trial can be expensive, which motivates asking how much improvement an optimizer can achieve within a small evaluation budget.

02

Physics, optimization, and software

NumPy and SciPy construct segment unitaries from a closed-system Hamiltonian and compose the two pulse segments in time order. An independent PennyLane implementation cross-checks the physics. A Gaussian process with a Matérn kernel selects candidates using expected improvement, after eight warm-start calls. The same counted objective interface constrains random search and bounded Powell. Runs persist configuration, seed, every evaluation, elapsed time, and status, supporting cancellation and exact replay.

03

The held-out experiment

The benchmark spans nine error settings: detuning −0.3, 0, or +0.3 crossed with amplitude error −0.1, 0, or +0.1. Three seeds give 27 runs per method. The target is average gate infidelity at or below 0.001 within a 40-call ceiling, including warm-start calls. The nominal pulse is evaluated once. The chart reports how many runs reached that target; the table provides final-error statistics so success counts are not the entire story.

04

What worked, and where

Random-start GP-EI reaches the target in 25 of 27 runs, with median final infidelity about 0.000252. A nominal-warm-start variant reaches it in all 27, with median about 0.000241; this variant was added after the first held-out run and is explicitly a post-hoc addition. Powell reaches the target in 9 of 27 runs and is much faster when detuning is zero. It stalls in the detuned settings, where GP-EI is more reliable. The result is conditional: Bayesian optimization helps with detuning, while Powell is excellent for the simpler amplitude-only problem.

05

Limits of the result

The simulator uses exact closed-system dynamics: no decoherence, measurement noise, or state-preparation and measurement error. There is one qubit and only two constant pulse segments. Calibration at one offset transfers poorly to another because the objective does not optimize robustness. Powell’s repeated seeds produce identical paths from the same nominal start. This is a controlled simulation study, not a hardware calibration claim or a demonstration on photonic devices.

Back to project categories ↑
Scientific & adaptive ML
THE FULL PROJECT · Adaptive language models

DriftGuard-GALA

A multilingual model-monitoring and adaptation experiment: when traffic shifts from English toward Hindi or Yoruba, detect the distribution change, choose an adaptation strategy, and measure whether the response actually improves the model.

Does retrieval help after a language shift?

Token F1 · higher is better · same scale for every language

Frozen baselineSelective retrieval
EnglishΔ +0.0000
0.3608
0.3608
HindiΔ +0.0276
0.2947
0.3223
YorubaΔ +0.0625
0.0254
0.0879
00.10.20.30.4 F1

English is unchanged. Hindi and Yoruba improve in this small recorded evaluation.

Only 2 prompts per language · 6 total · exploratory results
Read exact values in a table
LanguageBaselineRetrievalChange
English0.36080.36080.0000
Hindi0.29470.3223+0.0276
Yoruba0.02540.0879+0.0625
Token F1 from committed results. Higher is better; 2 prompts per language. Source ↗
01

The production question

A model can keep returning valid responses while the population of requests changes beneath it. DriftGuard replays an English-heavy baseline followed by shifted-language traffic, computes model-derived features, and uses Population Stability Index (PSI) to trigger an adaptation workflow. Drift detection is treated as the start of an investigation, not evidence that a particular update will help.

02

What was implemented

The model runtime uses Qwen 2.5 7B Instruct in four-bit quantization with bfloat16 compute and PEFT adapters. Experiments compare a frozen baseline, selective LoRA weight updates, replay-regularized variants, warmup and broader update variants, and retrieval-based in-context adaptation. The selective retrieval path preserves English behavior and adds retrieved exemplars for shifted Hindi and Yoruba prompts. The pipeline produces JSON metrics, reports, and visual comparisons.

03

How to read the chart

Token F1 measures overlap between generated and reference-answer tokens; higher is better. Each language now shows the frozen baseline and selective retrieval side by side on the same 0–0.4 scale. English remains 0.3608. Hindi rises from 0.2947 to 0.3223, and Yoruba rises from 0.0254 to 0.0879. The exact values and absolute changes are also shown in a table. Overall token F1 is approximately 0.2270 versus 0.2570.

04

The useful negative result

The online LoRA branches were implemented and tested, but the recorded ablations were harmful in this setup. The reported overall token-F1 change for those weight-update variants is −0.1979. Retrieval-based adaptation performed better here. This matters because the more complex intervention was not automatically the better operational choice: an adaptation system needs comparison against a frozen baseline and a way to reject unsuccessful updates.

05

How much confidence to place in it

The evaluation contains only two prompts per language—six in total. These results establish a working experiment and an exploratory comparison, not broad multilingual quality or statistically reliable superiority. A stronger follow-up would enlarge the held-out set, measure recovery over time, test repeated traffic shifts, and validate a gate that chooses among retrieval, no action, and separately trained adapters.

Back to project categories ↑
ML infrastructure
THE FULL PROJECT · Model infrastructure

ModelLaunch

A local model-release platform that treats deployment as a sequence of explicit decisions: validate the bundle, check model quality, deploy a candidate, test it, promote it, and keep watching after it becomes live. A failed candidate should leave the previous stable model available.

01Validate
02Preview
03Promote
04Monitor
One release. A way back.
Animated schematic of the local release workflow. Playback timing is illustrative, not a live deployment.
Local Kubernetes platform. Illustrated behavior from documented release drills. Source ↗
01

The engineering problem

A successful container start is not enough to declare an ML release healthy. The model may be inaccurate, its weights may be missing, requests may violate its schema, or latency may degrade only after promotion. ModelLaunch makes these cases observable and gives the release controller a defined recovery path.

02

The release contract

A versioned bundle contains an ONNX model, preprocessing configuration, request and response schemas, evaluation information, and a manifest with a SHA-256 payload hash. Quality thresholds are declared before a release and evaluated against a fixed release-validation set. A single-node kind cluster hosts blue/green deployments, preview and stable Services, Prometheus, and Grafana. FastAPI with ONNX Runtime exposes separate liveness, readiness, prediction, and metrics endpoints.

03

How promotion and rollback work

The candidate first passes integrity and quality validation, then preview correctness and load gates. The controller promotes the stable route only after those checks pass. A subsequent Prometheus verification window monitors latency and errors. Confirmed threshold breaches trigger rollback to the prior stable release. Each run writes a structured release record that can be replayed and inspected.

04

What the recorded drills show

The reference model released in 165 seconds with 12 gates passed. A degraded candidate with accuracy 0.425 failed the 0.9 gate before an image or pods were created. In the latency-fault drill, candidate p95 reached 868 ms against a 250 ms threshold; after two consecutive breaches, the previous model returned a healthy response 874 ms after detection. Detection itself took about 15.3 seconds after the fault was observed. A separate replica-kill drill recorded zero failures over 5,675 requests, and FlowBench released through the same gates in 155 seconds.

05

Scope and next steps

These are measured local experiments on the reported Apple M5 setup with two replicas, not availability guarantees for a production fleet. The distinction between fault detection time and recovery after detection matters. The project demonstrates release contracts, quality gates, health checks, observability, and incident recovery; broader cluster failure, multi-region operation, and organization-wide access controls require additional engineering.

Back to project categories ↑
ML infrastructure
THE FULL PROJECT · Developer tooling

ModelLedger

A command-line inventory and lifecycle checker for the model identifiers embedded in a codebase. ModelLedger turns those references into actionable findings and machine-readable output for continuous integration.

$ python -m modelledger.cli demo_repository

Inspect the bundled synthetic lifecycle records.

Synthetic demo registry. Critical retired-model findings block CI. Source ↗
01

The problem

A project can depend on a model that has been deprecated or retired even when its ordinary package dependencies remain healthy. ModelLedger scans source and configuration formats, discovers model references, and classifies them against lifecycle information so developers can see what needs attention.

02

What the tool produces

The scanner performs deterministic discovery and emits human-readable findings plus JSON output. Active, deprecated, and retired states distinguish routine inventory from migration work and critical failures. CI-friendly exit codes allow a pipeline to block on a retired-model finding rather than leaving the result in an unread log.

03

Demo and implementation scope

The bundled lifecycle registry uses clearly marked synthetic records. The portfolio’s scan interaction illustrates those records and their status handling; it does not scan a visitor’s files or claim that a named commercial model is actually retired. Separating discovery from lifecycle classification makes the output inspectable and keeps the provenance of lifecycle data explicit.

04

Why it belongs with ModelLaunch

The two tools address different parts of model operations: ModelLaunch controls how a candidate becomes a running service, while ModelLedger identifies model dependencies and lifecycle risks in the surrounding code. Together they show attention to deployment safety, maintenance, automation, and predictable command-line behavior.

Explore the implementation and evidence

Project repository ↗CLI and registry documentation ↗
Back to project categories ↑
Time-series & forecasting
THE FULL PROJECT · Quantitative research

Time-Series Momentum

A pre-registered replication of time-series momentum across 25 liquid ETFs, followed by an investigation into whether the performance survives leakage checks, transaction costs, multiple testing, and factor-exposure analysis. The central question is not simply whether the equity curve rises: it is what the evidence allows us to claim about that rise.

Original historical equity curve, 2000–2026, net of 5 basis points per side
Gross and net growth of $1. Log scale. The early ETF universe is narrower.
Original historical backtest figures, 2000–2026; net cost is 5 basis points per side. Source ↗
01

The strategy and research question

The primary specification follows the time-series momentum setup associated with Moskowitz, Ooi and Pedersen (2012): take the sign of trailing excess returns over a 12-month lookback, scale positions using volatility, and rebalance monthly. The designated configuration uses a 40% per-instrument volatility target. Signals are decided at time t and earn returns only through subsequent positions. The study uses ETF proxies across equities, fixed income, currencies, and commodities over 2000–2026.

02

What I built

The repository contains the signal and backtest engine, gross and net performance accounting, property-based leakage checks, expanding walk-forward validation, combinatorial purged cross-validation, multiple-testing corrections, and Fama–French five-factor plus momentum regressions. The hypothesis, configuration grid, and evaluation plan were committed before results. Raw prices are cached and random procedures seeded so the experiment can be repeated.

03

Testing whether the backtest deserves trust

Truncation invariance asks whether a signal changes when future data is removed; future poisoning asks whether corrupting future observations affects earlier decisions. Deliberately leaky positive controls verify that the tests can detect a problem. This process caught a full-sample quantile inside the volatility floor. Validation then uses overlap-based purging and a forward embargo instead of ordinary random folds. The nominal 12-configuration grid turned out to contain only six distinct strategies, because one volatility-target axis cancels under portfolio scaling; both counts are disclosed.

04

Performance and validation results

The primary historical backtest reports net Sharpe 0.706, annualized return 7.55%, annualized volatility 11.20%, and maximum drawdown −23.47%, after 5 basis points per side. A random-sign control through the same pipeline produces net Sharpe −0.045. Walk-forward out-of-sample Sharpe is 0.631; the five CPCV paths range from 0.546 to 0.706. The Deflated Sharpe Ratio passes across the tested effective-trial assumptions, but probability of backtest overfitting is 0.400. These are different diagnostics, not interchangeable seals of approval.

05

The finding beyond the headline

Passing multiple-testing corrections does not imply that returns are independent of known factors. UMD, the cross-sectional momentum factor, explains approximately 12% of monthly portfolio variation. Residual alpha is 3.97% per year, but its p-value moves from 0.0488 at Newey–West lag 21 to 0.0506 at lag 63. An initially striking currency-sleeve result also weakened after matching the analysis frequency to monthly rebalancing: the daily t-statistic of 7.31 becomes a monthly t-statistic of 2.84, and becomes insignificant after controlling for the other sleeves (t = 1.27, p = 0.20). The commodity sleeve’s independent loading remains an open research question.

06

Limitations and interpretation

ETFs are proxies for the futures instruments used in the original research, and the universe selects funds that still exist. Staggered inception means the early years are dominated by equities and Treasuries; the broadly diversified sample begins around 2007. Reported performance also declines across eras. The project supports a measured conclusion: the strategy has meaningful diversification and modest common momentum exposure, while the strength of residual-alpha evidence depends on statistical choices. It is a historical replication, not a promise of future returns.

Back to project categories ↑
Time-series & forecasting
THE FULL PROJECT · Adaptive forecasting

Adaptive Load Forecast

An online electricity-demand forecasting pipeline designed for a changing data stream. The system combines CAISO load with weather features and gives two learners different update speeds, then blends their predictions according to recent error.

EXPLORE THE LEARNING LOOP
CAISO loadWeather
↓ lag, calendar & weather features
↓ two predictions

The forecast weights each learner by its recent absolute errors. ADWIN monitors the error stream.

Architecture explorer; not a live forecast. Update cadence is configurable. Source ↗
01

Data and forecasting workflow

CAISO demand and Open-Meteo weather feed lag, calendar, and weather features. The model predicts demand, receives the observed outcome, and uses the resulting error to update the ensemble. The architecture explorer shows the roles of the learners and their combination; it is not a live electricity forecast.

02

Why two learning speeds

The fast branch is an online MLP that updates every step. The slower branch is an adaptive forest, updated every six steps by default. This configurable cadence lets the system combine rapid response with a second learning process. Recent absolute errors determine the blend, so a learner that has recently been less accurate contributes differently to subsequent predictions.

03

Detecting change

ADWIN monitors the error stream for change. Monitoring the forecast residual, rather than merely assuming the input distribution is stable, makes adaptation part of the forecasting loop. The project combines data ingestion, feature engineering, incremental models, and a serving simulator into an inspectable implementation.

04

Evidence and limits

The current portfolio presents the implemented architecture. A committed forecast trace is not available in the reviewed repository, so there is no invented observed-versus-predicted curve or accuracy improvement. The next evaluation should compare fixed and adaptive baselines on chronological held-out periods, report forecast error by horizon and operating regime, and inspect behavior around detected shifts.

Back to project categories ↑
Research implementations
THE FULL PROJECT · Alignment prototype

Online RLOO

A paper-inspired prototype of online REINFORCE with a leave-one-out baseline, organized around a fairer question: does updating a policy help beyond simply sampling several answers and selecting the best one?

Simulated policy
01Sample competing responses02Compute leave-one-out advantage03Update the simulated policy
Simulated policy; research prototype. Source ↗
01

Research basis and implemented scope

The repository cites Back to Basics: Revisiting REINFORCE-style Optimization for Learning from Human Feedback in LLMs, alongside the Aya Dataset paper for multilingual context. It implements the RLOO sampling, scoring, advantage calculation, update, and evaluation workflow. The current policy is simulated; this is an implementation of the experimental logic rather than a reproduction of the original paper’s large-model training results.

02

How the method works

For each prompt, sample four responses and score each one. A response’s advantage equals its reward minus the average reward of the other three responses. That leave-one-out comparison assigns credit relative to competing outputs for the same prompt. Training uses English and Hindi prompts; evaluation includes Yoruba as a held-out language. Generation, reward scoring, alignment, and evaluation are separate modules.

03

The comparison and recorded outcome

The experiment compares a frozen one-response baseline, best-of-four sampling without updates, and online RLOO. Reported mean rewards are 0.3864, 0.7463, and 0.8523 respectively; win rates are 0.2500, 0.5833, and 0.7500. The best-of-four control is essential: it asks whether the apparent improvement comes from policy changes or simply additional sampling. These numbers describe the simulated MVP and its reward mechanism.

04

What remains to be reproduced

A real QLoRA training loop, a stronger multilingual reward model, larger Aya evaluation slices, actual token and latency accounting, and comparison against SFT or DPO remain future work. The prototype demonstrates orchestration and estimator logic; it does not yet demonstrate large-model alignment gains or reliable cross-lingual transfer.

Back to project categories ↑
Research implementations
THE FULL PROJECT · Reasoning-budget prototype

Budget-Forced Test-Time Scaling

A research prototype exploring how a frozen model’s inference policy can allocate reasoning effort: let easy tasks stop early, give harder tasks room to continue, and enforce an upper budget.

Simulated decoding workflow
01Estimate task difficulty02Assign a reasoning budget03Continue or stop generation
Simulated decoding workflow; no live model calls. Source ↗
01

Ideas from the papers

The repository draws on s1: Simple test-time scaling for minimum-budget forcing and forced continuation, and Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models for difficulty-aware allocation. It also discusses work on uncertainty verbalization and the effect of removing thinking tokens. The implemented contribution combines these ideas into a configurable generation-control experiment.

02

The three policies

The baseline accepts the model’s natural stopping point. Minimum-budget forcing suppresses an early answer transition and injects a Wait-style continuation until enough reasoning budget is consumed. The plan-and-budget policy assigns minimum and maximum budgets according to estimated task difficulty. The runtime tracks forced continuation, forced termination, token use, answer accuracy, and overthinking.

03

Recorded prototype results

At the reported 256-token ceiling, minimum-budget forcing and plan-and-budget both reach accuracy 1.0000 in the compact simulated benchmark. Average total token counts are 264 and 152 respectively, and recorded overthinking rates are 0.3333 and 0.0000. The ceiling applies to reasoning; reported total tokens include the answer. These outcomes illustrate the policy trade-off within the simulator, not measured performance from a deployed reasoning LLM.

04

Replication boundary

The current runtime simulates generation. It does not reproduce the papers’ full training, models, or benchmark-scale evaluations. Next steps are a real Transformers or vLLM wrapper, model-specific delimiter control, MATH or AIME evaluation subsets, and measured inference cost and latency. The engineering question is whether adaptive budget control retains answer quality while reducing unnecessary computation.

Back to project categories ↑
Research implementations
THE FULL PROJECT · Data-selection prototype

Malleable-Eval

A data-selection and evaluation prototype that treats the next training examples as a decision: use a limited teacher-labeling budget on prompts where the student appears least reliable, then compare that choice with random sampling at the same budget.

Proxy student · deterministic teacher
01Find informative prompts02Relabel and validate examples03Compare against random selection
Proxy student and deterministic teacher; research prototype. Source ↗
01

The idea and implementation

The pipeline runs a student over a compact Aya-style multilingual pool, ranks prompts using uncertainty, rarity, and prompt length, sends selected examples through a teacher policy, validates the resulting pairs, and retrains the student. It contains separate orchestration, data-validation, and evaluation modules with automatically generated reports. The aim is a repeatable experiment in adaptive dataset construction.

02

A same-budget comparison

The recorded run begins with six training examples. Both random selection and targeted top-k selection expand that set to ten. Token F1 is 0.0794 for the seed-only baseline, 0.0863 with random additions, and 0.0879 with targeted additions. Exact match remains zero, and the other reported metrics do not distinguish the two expanded sets. The small token-overlap gain is exploratory rather than a broad performance claim.

03

Research scope

The current student is a lightweight memory-based proxy and the teacher is deterministic. The selection score approximates informativeness; it is not actual checkpoint-level Variance of Gradients. This should be presented as a research-inspired workflow prototype, not a completed reproduction of a transformer-training paper. The repository documents the gap between the implemented MVP and its intended research machinery.

04

What would make the next experiment stronger

Replace the proxy with a real model, compute gradient-variance signals across checkpoints, connect a capable teacher, and evaluate on a larger held-out corpus with repeated budgets and seeds. Validation before training remains central: selecting interesting prompts is useful only if the synthesized examples are valid and the evaluation can distinguish improvement from noise.

Back to project categories ↑
Research figure

Original figure at readable size. Scroll horizontally on a small screen.

A LITTLE ABOUT ME

I like the part where
an idea becomes real.

I’m Abdul Ali Mamnun, a data scientist. I build projects around machine learning, scientific computing, and the systems that make models useful.

My work ranges from neural fluid prediction and local model deployment to forecasting, multilingual adaptation, and careful backtesting. This collection brings the code and the experiments together.

Find me on GitHub ↗