Minimize simulated inventory cost by tuning four replenishment controls within 48 noisy queries
Optimization, Planning & Controlsimulation optimizationinventory control
Background
Simulation optimization designs search algorithms for systems whose objective can only be sampled, noisily, from a stochastic simulator, where the standing difficulty is doing well within a few dozen costly runs. The starting point is a uniform-random search routine that proposes four normalized controls of a replenishment policy and receives noisy cost feedback. The work is to redesign that routine into a query-efficient search. The result must hold across unseen demand and cost regimes, where noise alone can make a mediocre setting look best.
instruction.mdthis is what the agent is given
Design a reproducible optimizer for four normalized controls of a replenishment policy under stochastic demand and lead times. Minimize simulated operating cost within 48 queries over [0, 1]^4. You inherit a uniform-random starter and receive noisy objective feedback through a bounded ask/tell interface.
Hard Constraints
Edit only /app/methods/main/; solver.py must define Optimizer.
Use only the Python standard library and NumPy 2.2.6.
The constructor is Optimizer(dim, lower, upper, budget, seed, rng).
ask(n) must return a finite matrix with 1 through n rows, exactly dim columns, and every coordinate inside the supplied bounds.
The trusted evaluator owns the simulator, independent rescoring, query counter, and timeout. Extra returned points do not increase the query budget.
Import failures, crashes, malformed output, non-finite values, and out-of-bounds proposals invalidate the submission. If the aggregate runtime expires, completed valid work is retained under the published partial-work rule.
The submitted process cannot read or modify trusted evaluator assets and has no verifier network access.
The sealed evaluator runs many short optimizer sessions under one aggregate runtime boundary. Keep each ask and tell bounded and vectorized; expensive dense refits or very large candidate scans at every query can exhaust that shared budget.
What You Have
/app/data/visible.json contains public development cases spanning the same inventory-policy regimes as evaluation.
/app/data/visible_anchors.json contains public calibration traces used by the matched visible scorer.
/app/methods/main/solver.py is the uniform-random weak starter.
/app/selfcheck.py runs the same query, independent-rescoring, aggregation, and mapping semantics on public cases.
The four coordinates encode reorder level, order-up-to level, expedite threshold, and smoothing, each normalized to [0, 1].
The supplied rng is np.random.default_rng(seed) and should drive optimizer randomness. Simulator noise is controlled by the trusted evaluator.
The evaluator calls ask(n), evaluates the returned controls, and calls tell(X, y) or tell(X, y, metadata) when accepted. Values in y are costs, so lower is better.
A positive integer self.batch may request a preferred batch size; the evaluator caps it to the remaining budget.
What You Submit
Submit general optimizer code, not a final control vector or precomputed case answers. The submission must be self-contained in /app/methods/main/solver.py; sibling modules are not copied to the trusted verifier.
How It Is Judged
The trusted parent re-runs the optimizer on distinct cases with the same 48-query interface used publicly. When an observed incumbent improves, it is independently re-simulated with a candidate-keyed random stream. Best-so-far independent traces are summarized across repeated runs, with both sustained progress and final quality contributing to the metric. Better cost reduction across the full case mixture is better. Calibration assets and the exact leaderboard transformation remain verifier-only.
Rollouts
83 minWall clock
$28.16Spend
40.5MTokens
61Versions, 11 kept
On the visible set
keptrolled backsubmitted
v0The agent started from the shipped uniform-random baseline01 min · $0.24
v1The agent swapped random sampling for an even space-filling sweep0.0869275 min · $0.65
v2The agent opened with hand-built policies spanning the four regimes0.2981026 min · $0.90
v3The agent tried mutating its best observations, and one case got all the gain0.2981768 min · $1.18
v4The agent tried a Gaussian process, and it bought little for double the runtime0.29814510 min · $1.48
v5The agent tried a systematic axis sweep, and one lean case collapsed0.28403114 min · $2.13
v6The agent read early costs to tell lean cases from service cases0.30336416 min · $2.42
v7The agent tried better-looking probes, and noise kept them off the record0.29710418 min · $2.78
v8The agent shortened its probing and spent the rest refining near its best0.30355120 min · $3.10
v9The agent tried jittering only the expedite knob, and all-axis moves won0.30354422 min · $3.46
v10The agent tried refining only its single best point, and reserved cases fell0.3039124 min · $3.90
v11The agent tried halving its step size, and the moves became redundant0.30377424 min · $4.06
v12The agent doubled its step size and the gain repeated across noise panels0.30392825 min · $4.33
v13The agent tried stepping even wider, and it degraded its own best policies0.30373126 min · $4.53
v14The agent tried cutting the lean branch short, and one case turned unstable0.30367427 min · $4.76
v15The agent tried probing longer before refining, and the balance got worse0.30370927 min · $4.98
v16The agent routed on the whole cost trend instead of just its endpoints0.30399230 min · $5.48
v17The agent tried keeping two parents, and the third proved a useful hedge0.30392930 min · $5.71
v18The agent tried flattening its parent weights, and the gain did not repeat0.30383631 min · $6.01
v19The agent tried steering moves per regime, and one panel collapsed0.30399633 min · $6.33
v21The agent tried searching in policy coordinates, and lost local diversity0.30383334 min · $6.86
v22The agent tried adding a stronger opening policy, and record timing broke0.29568336 min · $7.22
v23The agent tried replacing its first probe, and that probe was essential0.30297936 min · $7.48
v24The agent tried averaging with neighbours to fight the winner's curse0.30383338 min · $8.25
v25The agent tried forcing its parents apart, and one basin can deserve several0.30390338 min · $8.51
v26The agent tried mixing several step scales, and one fixed scale was steadier0.30368739 min · $8.78
v27The agent stopped shrinking its step late and kept refining to the end0.30400540 min · $9.12
v28The agent tried a smaller fixed step, and it clearly lost0.3037541 min · $9.38
v29The agent tried a larger fixed step, and it also lost0.30390341 min · $9.65
v30The agent added a second lean policy near the holding optimum0.3040342 min · $10.10
v31The agent tried reordering the lean probes, and the old order hedged betteron a subset: 0.02410644 min · $10.79
v32The agent tried a simpler routing rule, and it lost narrowly but always0.30410746 min · $11.40
v33The agent tried zeroing the expedite probe, and lost a needed contrast0.29726946 min · $11.63
v34The agent retested single-parent refinement, and hedging still won0.30385647 min · $11.95
v35The agent tried a fourth parent, and three kept the better balanceon a subset: 0.0243448 min · $12.33
v36The agent tried broad parents early and narrow late, and validation disagreedon a subset: 0.02440349 min · $12.81
v37The agent tried mirrored pairs of moves, and updating each query was better0.30397150 min · $13.13
v38The agent tried moving one control at a time, and joint moves were stabler0.30292951 min · $13.53
v39The agent tried cautious routing, and lost more lean cases than it saved0.30386253 min · $14.12
v40The agent tried reordering only for confident lean cases, and gradual wonon a subset: 0.02417555 min · $14.68
v41The agent tried interleaving near-duplicate resamples, and lost refinements0.30386555 min · $15.02
v42The agent tried exploiting its best parent harder, and the old mix was safer0.30405757 min · $15.54
v43The agent tried steering along elite differences, and random moves held up0.30396357 min · $15.89
v44The agent tried a fast path for expensive-looking cases, and it fired too often0.30330860 min · $16.73
v45The agent tightened that fast path, and the reserved cases still lost0.30404561 min · $17.96
v46The agent tightened the fast path once more, then dropped the branchon a subset: 0.02432762 min · $18.42
v47The agent tried a third branch for service-like cases, and routing noise hurt0.30376863 min · $19.06
v48The agent tried tiny jitter to farm extra scoring draws, and quality collapsed0.29438365 min · $20.16
v49The agent tried refining the centroid of its best three, and blurred them0.30389367 min · $20.70
v50The agent tried swapping probes for costly-looking sessions, and it did not last0.30392468 min · $21.19
v51The agent tried adding service variants mid-run, and they crowded out refining0.303969 min · $21.60
v52The agent tried pinning proposals to the box edges, and inward moves mattered0.30385670 min · $22.01
v53The agent tried uniform instead of Gaussian moves, and missed the tails0.30394970 min · $22.41
v54The agent tried a smaller step for lean cases, and they liked the shared one0.30392471 min · $22.82
v55The agent tried a larger step for lean cases, and stayed below the incumbent0.30397772 min · $23.22
v56The agent tuned the lean-only step once more, and the shared step still wonon a subset: 0.02426573 min · $23.74
v57The agent tried a larger step for non-lean cases, and it also lost0.30384173 min · $24.15
v58The agent swept its routing thresholds, and the winner failed fresh panels0.30407677 min · $25.14
v59The agent tried following the direction of past wins, and noise gave none0.30384778 min · $25.62
v60The agent asked for its fixed probes in one batch to cut round trips0.30366779 min · $26.51
v61The agent batched the branch probes too and fuzz-tested the whole interface0.30366780 min · $26.96
On the hidden set
Case
Starter · 0.0
Reference · 0.6
Upper · 1.0
This run (GPT-5.6-sol)
Case 01
0.000000
0.040052
1.000000
0.046930
Case 02
0.000000
0.041445
1.000000
0.050016
Case 03
0.000000
0.008948
1.000000
0.008686
Case 04
0.000000
0.009260
1.000000
0.028073
Case 05
0.000000
0.026032
1.000000
0.046464
Case 06
0.000000
0.020090
1.000000
0.029066
Case 07
0.000000
0.031189
1.000000
0.028998
Case 08
0.000000
0.020216
1.000000
0.016136
Case 09
0.000000
0.018509
1.000000
0.014298
Case 10
0.000000
0.026902
1.000000
0.029309
Case 11
0.000000
0.016185
1.000000
0.024458
Case 12
0.000000
0.016510
1.000000
0.015960
Case 13
0.000000
0.029049
1.000000
0.034305
Case 14
0.000000
0.017305
1.000000
0.025874
Case 15
0.000000
0.007159
1.000000
0.015928
Case 16
0.000000
0.011258
1.000000
0.007920
Normalised score
0.5484
376 minWall clock
$41.46Spend
67.0MTokens
8Versions, 7 kept
On the visible set
keptrolled backsubmitted
v0The agent inherited the shipped uniform-random search baseline034 min · $3.25
v1The agent added a kernel-ridge surrogate with trust region and informed design0.2788235 min · $3.43
v2The agent shortened the informed design to six best-prior-first points0.28934 min · $3.25
v3The agent cut the design to three points and retuned the trust region0.30356 min · $5.79
v4The agent rewrote the surrogate as a low-rank family basis0.2944141 min · $16.93
v5The agent shortened screening and switched to a lower-confidence-bound proposal0.2998200 min · $21.71
v6The agent replaced the Gaussian posterior with a nonparametric particle cloud0.3022253 min · $27.11
v7The agent rebuilt the basis and cloud from 4200 broader instances0.3025375 min · $41.23
On the hidden set
Case
Starter · 0.0
Reference · 0.6
Upper · 1.0
This run (Opus 5)
Case 01
0.000000
0.040052
1.000000
0.051208
Case 02
0.000000
0.041445
1.000000
0.032326
Case 03
0.000000
0.008948
1.000000
0.011652
Case 04
0.000000
0.009260
1.000000
0.028254
Case 05
0.000000
0.026032
1.000000
0.040151
Case 06
0.000000
0.020090
1.000000
0.027992
Case 07
0.000000
0.031189
1.000000
0.031041
Case 08
0.000000
0.020216
1.000000
0.012016
Case 09
0.000000
0.018509
1.000000
0.024271
Case 10
0.000000
0.026902
1.000000
0.023772
Case 11
0.000000
0.016185
1.000000
0.021966
Case 12
0.000000
0.016510
1.000000
0.015480
Case 13
0.000000
0.029049
1.000000
0.034983
Case 14
0.000000
0.017305
1.000000
0.030843
Case 15
0.000000
0.007159
1.000000
0.023293
Case 16
0.000000
0.011258
1.000000
0.016949
Normalised score
0.5544
15 minWall clock
$1.46Spend
6.1MTokens
5Versions, 4 kept
On the visible set
keptrolled backsubmitted
v0The agent inherited the shipped uniform-random search baseline0
v1The agent fitted a basic GP with fixed lengthscales and LCB0.125176
v2The agent switched to TuRBO-GP with rank warping and simplex polishing0.902636
v3The agent added multi-band trust regions and quadratic refinement0.90656
v4The agent added a dual incumbent and batched ask support0.907474
On the hidden set
Case
Starter · 0.0
Reference · 0.6
Upper · 1.0
This run (Gemini 3.7 Flash)
Case 01
0.000000
0.040052
1.000000
0.045184
Case 02
0.000000
0.041445
1.000000
0.022160
Case 03
0.000000
0.008948
1.000000
0.012565
Case 04
0.000000
0.009260
1.000000
0.002686
Case 05
0.000000
0.026032
1.000000
0.006561
Case 06
0.000000
0.020090
1.000000
0.003574
Case 07
0.000000
0.031189
1.000000
0.017922
Case 08
0.000000
0.020216
1.000000
0.000000
Case 09
0.000000
0.018509
1.000000
0.000866
Case 10
0.000000
0.026902
1.000000
0.011020
Case 11
0.000000
0.016185
1.000000
0.017712
Case 12
0.000000
0.016510
1.000000
0.000046
Case 13
0.000000
0.029049
1.000000
0.012422
Case 14
0.000000
0.017305
1.000000
0.006438
Case 15
0.000000
0.007159
1.000000
0.000000
Case 16
0.000000
0.011258
1.000000
0.000000
Normalised score
0.1883
351 minWall clock
$21.91Spend
54.6MTokens
11Versions, 10 kept
On the visible set
keptrolled backsubmitted
v0The agent inherited the shipped uniform-random search baseline0$0.54
v1The agent built a GP-BO with Latin-hypercube init and expected improvement0.199$1.08
v2The agent dropped output standardization so the GP interpolated raw costs0.234$2.61
v3The agent added an inhibition radius and multi-scale local candidates0.244$3.58
v4The agent replaced the init with six regime-spanning warm-start points0.279$4.16
v4_tmpThe agent snapshotted the wrong code as v4_tmp and deleted itmis-snapshot, deleted$4.69
v5The agent made candidates local-heavy and lowered the assumed noise0.287$5.22
v6The agent widened the random candidate pool and raised the lengthscale0.2972$6.91
v7The agent rewrote the solver cleanly on a Matern-3/2 kernel0.2966$8.65
v8The agent lowered inhibition and lengthened the final re-evaluation tail0.2945$9.81
v9The agent added a bounds-safety clip to its proposals0.2945$16.61
On the hidden set
Case
Starter · 0.0
Reference · 0.6
Upper · 1.0
This run (Kimi K3)
Case 01
0.000000
0.040052
1.000000
0.062841
Case 02
0.000000
0.041445
1.000000
0.030342
Case 03
0.000000
0.008948
1.000000
0.008403
Case 04
0.000000
0.009260
1.000000
0.027588
Case 05
0.000000
0.026032
1.000000
0.040941
Case 06
0.000000
0.020090
1.000000
0.021597
Case 07
0.000000
0.031189
1.000000
0.030523
Case 08
0.000000
0.020216
1.000000
0.013728
Case 09
0.000000
0.018509
1.000000
0.018959
Case 10
0.000000
0.026902
1.000000
0.026911
Case 11
0.000000
0.016185
1.000000
0.021426
Case 12
0.000000
0.016510
1.000000
0.017365
Case 13
0.000000
0.029049
1.000000
0.035927
Case 14
0.000000
0.017305
1.000000
0.023651
Case 15
0.000000
0.007159
1.000000
0.015185
Case 16
0.000000
0.011258
1.000000
0.014535
Normalised score
0.5618
38 minWall clock
$10.33Spend
15.4MTokens
20Versions, 8 kept
On the visible set
keptrolled backsubmitted
v0The agent inherited the shipped uniform-random search baseline0
v1The agent added a designed (s,S) prototype start with compass and GP-LCB0.02336
v2The agent lengthened the design and added confirmation re-evals and line searches0.0229
v3The agent added mean-best repeats and a frozen-axis ridge refinement0.02339
v4The agent widened the design to twenty anytime ridge-cover points0.02394
v5The agent moved holding and service prototypes earlier and searched on-face0.02572
v6The agent replaced the lucky first point with a true-mean holding policy0.292
v7The agent probed more early main-face points and drew unlucky scores0.287
v8The agent swapped in new prototypes and lost the holding complement0.295
v9The agent kept the v5 prefix and added four complements plus face fill0.30471
v10The agent capped the plan at 24 points and made the grid adaptive0.30468
v11The agent mixed all three faces into the adaptive grid0.30465
v12The agent re-queried the top five mean-best unique points0.30463
v13The agent swapped one holding plan point for a truly better one0.30463
v14The agent swapped a service plan point so a lucky miss cannot block0.30503
v15The agent also swapped a second plan slot for no gain0.30503
v16The agent replaced a duplicate extra point without touching the prefix0.30506
v17The agent replaced the golden face fill with GP-LCB candidates0.30505
v18The agent swapped two late plan slots for no gain0.30506
v19The agent jittered the late face fill0.30499
On the hidden set
Case
Starter · 0.0
Reference · 0.6
Upper · 1.0
This run (Grok 4.6)
Case 01
0.000000
0.040052
1.000000
0.046317
Case 02
0.000000
0.041445
1.000000
0.051758
Case 03
0.000000
0.008948
1.000000
0.016195
Case 04
0.000000
0.009260
1.000000
0.020208
Case 05
0.000000
0.026032
1.000000
0.029560
Case 06
0.000000
0.020090
1.000000
0.028154
Case 07
0.000000
0.031189
1.000000
0.030137
Case 08
0.000000
0.020216
1.000000
0.016462
Case 09
0.000000
0.018509
1.000000
0.024936
Case 10
0.000000
0.026902
1.000000
0.027918
Case 11
0.000000
0.016185
1.000000
0.016650
Case 12
0.000000
0.016510
1.000000
0.023480
Case 13
0.000000
0.029049
1.000000
0.030933
Case 14
0.000000
0.017305
1.000000
0.016726
Case 15
0.000000
0.007159
1.000000
0.025000
Case 16
0.000000
0.011258
1.000000
0.002537
Normalised score
0.5524
69 minWall clock
$1.32Spend
11.9MTokens
15Versions, 9 kept
On the visible set
keptrolled backsubmitted
v0The agent inherited the shipped uniform-random search baseline0$0.33
v1The agent added an informed design with a quadratic surrogate0.193540 min · $0.66
v2The agent switched to a two-basin design with Gaussian sampling0.2692$0.72
v3The agent learned the design greedily from observation means over synthetic cases0.2655$0.78
v4The agent added a quadratic-surrogate center and a final re-evaluation phase0.2672$0.83
v5The agent centered sampling on exponentially weighted means with shrinking sigma0.2851$0.89
v6The agent embedded a ridge predictor trained on synthetic instances0.2829$0.94
v7The agent split sampling into two basin centers with softmax allocation0.29758 min · $1.00
v7_anisoThe agent widened one axis around a single noisy center0.2787$1.04
v7_threecThe agent split the search into three basin centers0.2965$1.08
v7_twocThe agent introduced per-basin argmin centers with anisotropic sigma0.2928$1.12
v7_twoc sweepThe agent swept the basin allocation temperature and picked 1000.2961$1.16
v7_twocwThe agent weighted the basin centers by region instead of argmin0.2874$1.20
v8The agent added a tight final sampling phase and robustness guards0.29700867 min · $1.25
v8-v10The agent tried merges, re-evals and pooled centers, all plateauing0.277$1.28
On the hidden set
Case
Starter · 0.0
Reference · 0.6
Upper · 1.0
This run (DeepSeek V4 Pro)
Case 01
0.000000
0.040052
1.000000
0.049147
Case 02
0.000000
0.041445
1.000000
0.050940
Case 03
0.000000
0.008948
1.000000
0.017324
Case 04
0.000000
0.009260
1.000000
0.029638
Case 05
0.000000
0.026032
1.000000
0.042477
Case 06
0.000000
0.020090
1.000000
0.025403
Case 07
0.000000
0.031189
1.000000
0.025441
Case 08
0.000000
0.020216
1.000000
0.017100
Case 09
0.000000
0.018509
1.000000
0.024357
Case 10
0.000000
0.026902
1.000000
0.043293
Case 11
0.000000
0.016185
1.000000
0.017453
Case 12
0.000000
0.016510
1.000000
0.013942
Case 13
0.000000
0.029049
1.000000
0.029090
Case 14
0.000000
0.017305
1.000000
0.028130
Case 15
0.000000
0.007159
1.000000
0.017049
Case 16
0.000000
0.011258
1.000000
0.012194
Normalised score
0.5704
107 minWall clock
$7.57Spend
24.1MTokens
25Versions, 9 kept
On the visible set
keptrolled backsubmitted
v1The agent inherited the shipped uniform-random search baseline03 min · $0.15
v2-v4The agent iterated GP-BO builds on five-case subsets without snapshotting5-case subsets only$1.04
v5The agent consolidated the GP-BO with a fixed eight-point designsubset 0.0143-0.049435 min · $1.92