# Profile evaluation campaign

Inventory: **10 model repos, 11 Spaces and 9 datasets**. Eight model repos contain runnable artifacts; two are documents. **No external rank 1 is verified.**

Native readback covers 48 official endpoints and 1,021 returned entries. None match these model IDs. A field-matching mistake in the first derived readback is corrected; original responses remain preserved. No pagination-completeness claim.

Existing failures, used confirmation and eight sealed well groups remain preserved. The new adapter confirmation is now consumed and may not be retuned or called fresh again.

## Models

| Artifact | Actual task | Current evidence / evaluation |
|---|---|---|
| [sankalpsthakur](https://huggingface.co/sankalpsthakur/sankalpsthakur) | Personal profile document | No model weights or input contract |
| [tiny-industrial-edge-onnx](https://huggingface.co/sankalpsthakur/tiny-industrial-edge-onnx) | Synthetic six-channel tabular classification | Four authored ONNX rows executed |
| [unmute-voice-stack-card](https://huggingface.co/sankalpsthakur/unmute-voice-stack-card) | Voice stack architecture document | No STT/TTS model or running service |
| [forge-tiny-drift-multiruntime](https://huggingface.co/sankalpsthakur/forge-tiny-drift-multiruntime) | Fifty-cycle drift classification | Sixteen authored NumPy/ONNX windows pass 1e-4 probability and exact decision parity |
| [forge-electrolyser-telemetry-baseline](https://huggingface.co/sankalpsthakur/forge-electrolyser-telemetry-baseline) | Causal telemetry residual scoring | Seventy-three authored windows scored; exact batch-prefix comparison fails at 1.33e-15; rowwise prefix and independent equation checked |
| [forge-pump-surrogate-multiruntime](https://huggingface.co/sankalpsthakur/forge-pump-surrogate-multiruntime) | Six-output synthetic pump surrogate | All eight public frozen vectors pass the original 1e-3 PyTorch/ONNX gate; the card reports an older 32-vector receipt |
| [qwen3-06b-typed-decisions-cloud-pilot](https://huggingface.co/sankalpsthakur/qwen3-06b-typed-decisions-cloud-pilot) | Grounded Yes/No language decisions | One completed included CPU attempt: 128 unused BoolQ and 32 unused grounded StrategyQA; fixed base/adapter and temperatures |
| [industrial-edge-setpoint-kernels](https://huggingface.co/sankalpsthakur/industrial-edge-setpoint-kernels) | Compact supervisory dynamics and control research | Both pinned linear/RFF predictors execute three authored state-command pairs and reject malformed/nonfinite inputs; prior PI/MPC losses retained |
| [well-production-event-causal](https://huggingface.co/sankalpsthakur/well-production-event-causal) | Causal production-event detection | Eight numeric checkpoints reload on 128 eligible probes each; three stateful references replay 80,534 old confirmation decisions; reserved groups untouched |
| [controlled-diet-ogi-mask-reference](https://huggingface.co/sankalpsthakur/controlled-diet-ogi-mask-reference) | Exploratory generated-mask gas-imaging reference | Seven frozen arms replay 21 training-only grids at original 1e-10/1e-9 tolerance with exact threshold counts; both learned arms still lose the spatial prior |

## Spaces

| Artifact | Actual task | Current evidence / evaluation |
|---|---|---|
| [transient-energy-ems](https://huggingface.co/spaces/sankalpsthakur/transient-energy-ems) | Fictional energy workflow and mock connectors | Validate data provenance and actual optimization before any energy-saving claim |
| [sam-fruit-what-does-sam-add](https://huggingface.co/spaces/sankalpsthakur/sam-fruit-what-does-sam-add) | Saved segmentation teaching ablation, no live SAM call | Bind a licensed physical image dataset, real segmentation checkpoint and object-disjoint evaluator before claiming model performance |
| [games-strategy-demo](https://huggingface.co/spaces/sankalpsthakur/games-strategy-demo) | Saved synthetic simplified-game learner | Use an independently specified executable game, frozen opponents, equal interaction budgets and fresh episodes before strategic ranking |
| [forge-industrial-agent-lab](https://huggingface.co/spaces/sankalpsthakur/forge-industrial-agent-lab) | Offline engineering worksheets and inspection/audit interfaces | Calculation/admission correctness, refusal and stale-output behavior; upstream PP-OCRv5 has nine exact reads among eleven authored crops, two retained wrong reads and no physical meter qualification |
| [game-theory-lab](https://huggingface.co/spaces/sankalpsthakur/game-theory-lab) | Educational payoff calculations | Test mathematical solutions and input refusal; no universal model leaderboard |
| [qaoa-depth-noise-lab](https://huggingface.co/spaces/sankalpsthakur/qaoa-depth-noise-lab) | Saved quantum/classical result explorer | Reproduce the declared instance/circuit/noise budget and compare strong classical references; do not turn saved evidence into a live QPU claim |
| [forge-pump-edge-twin-lab](https://huggingface.co/spaces/sankalpsthakur/forge-pump-edge-twin-lab) | Actual browser ONNX pump inference | Model replay, input units/envelope and stale-output clearing; browser availability is separate from prediction accuracy |
| [typed-decisions-qwen3-demo](https://huggingface.co/spaces/sankalpsthakur/typed-decisions-qwen3-demo) | Live included ZeroGPU grounded-language endpoint | Verify actual requests, serving revision, posterior/calibration presentation and input refusal; use a separate frozen evaluator for accuracy |
| [industrial-edge-setpoint-lab](https://huggingface.co/spaces/sankalpsthakur/industrial-edge-setpoint-lab) | Saved control-research explorer | Check all controller/scenario selections against frozen results and make PI/MPC losses visible |
| [well-event-transfer-audit](https://huggingface.co/spaces/sankalpsthakur/well-event-transfer-audit) | Saved well/foundation/incremental research explorers | Check every candidate/group against pinned results and show alarm-budget, censoring, eligibility, timing and portability failures |
| [cd-ogi-mask-reference-audit](https://huggingface.co/spaces/sankalpsthakur/cd-ogi-mask-reference-audit) | Saved laboratory gas-mask explorer | Check saved comparisons and data modality limits; it performs no live leak inference |

## Datasets

| Artifact | Actual task | Current evidence / evaluation |
|---|---|---|
| [iea-global-ev-outlook-2026-owid-processed](https://huggingface.co/datasets/sankalpsthakur/iea-global-ev-outlook-2026-owid-processed) | Processed aggregate energy data | Data provenance and denominator validation, no predictive model rank |
| [banking77-gemma-evaluation-receipt](https://huggingface.co/datasets/sankalpsthakur/banking77-gemma-evaluation-receipt) | Saved Banking77 evaluation receipt | No learned checkpoint in this dataset; retain untouched-test and local OOM/failure limits |
| [forge-industrial-control-scenarios](https://huggingface.co/datasets/sankalpsthakur/forge-industrial-control-scenarios) | Synthetic generic control scenarios | Simulation comparator inputs, not measured plant truth |
| [ibm-150q-ising-hardware-results](https://huggingface.co/datasets/sankalpsthakur/ibm-150q-ising-hardware-results) | Saved hardware/quantum audit results | Immutable circuit/shot/instance receipts and classical comparison, no live model endpoint |
| [forge-pump-digital-twin-synthetic](https://huggingface.co/datasets/sankalpsthakur/forge-pump-digital-twin-synthetic) | Clean-room synthetic pump data | Generator approximation benchmark, not physical pump validation |
| [johansson-delayed-sensor-control](https://huggingface.co/datasets/sankalpsthakur/johansson-delayed-sensor-control) | Synthetic checked quadruple-tank data | Simulation-only causal observer/control protocol |
| [3w-causal-real-group-audit](https://huggingface.co/datasets/sankalpsthakur/3w-causal-real-group-audit) | Real-source well-study audit rows | Audit tables, not redistributed physical source recordings or a native benchmark |
| [cd-ogi-recording-mask-audit](https://huggingface.co/datasets/sankalpsthakur/cd-ogi-recording-mask-audit) | Generated-mask recording audit rows | Audit rows, not images, independent physical trials or chemical pixel truth |
| [tadi-release-denominator-audit](https://huggingface.co/datasets/sankalpsthakur/tadi-release-denominator-audit) | Gas-release/session denominator audit | Qualified source metadata scope only; metric-specific physical validation remains separate |

## Next measured work

The fresh adapter result is mixed: BoolQ 107/128 versus base 99/128; grounded StrategyQA 18/32 versus base 19/32. Both record-bootstrap accuracy intervals include no improvement. Raw confidence worsens; fixed temperatures do not establish cross-source calibration. Five of sixteen older GPU/CPU probe comparisons exceed the unchanged 0.002 tolerance. No checkpoint is promoted.

For each predictive track, the JSON disposition specifies the stronger reference, data grouping, eligibility and failure metrics needed before another fit. Most industrial tasks currently lack a registered external leaderboard and independent physical qualification. Native score files will only follow a complete conforming run on an eligible benchmark.

Use free/included compute. Included GPU refresh was observed for October 10 at 00:00 UTC (04:00 Dubai); CPU execution proceeded now. TPU compatibility is unqualified. Paid compute, plant/robot access, new human outreach and profile edits remain separate authorization boundaries.
