How we assess the real, origin-level price of physical commodities — the data we use, how we combine it, how we reconstruct landed cost, and how we approach forecasting. Written to be checked, and honest about what it can and can’t do.
The Commtora Index (CTX) is an indicative, assessed benchmark for the real, origin-level price of the physical commodities we trade. It sits in the tradition of physical-market price assessments(as Platts or Argus do for their markets) — a trade reference, not a financial or regulated benchmark.
It is not a licensed exchange or futures value; not a settlement benchmark for financial instruments; not a firm, executable quote (that comes from the desk); and not investment advice. Wherever a number is modelled or estimated rather than observed, we label it as such.
Every figure is tagged by source and confidence. We never invent data to fill a gap — a stale point is flagged, not faked.
We combine sources with a robust central tendency (median / winsorized mean) so a single stray print cannot move the mark, weighted by recency and source reliability.
CTX is published as a band, not a single point. A false-precision number would misrepresent what we know; the width of the band reflects source dispersion and staleness. When fresh data is absent we carry the last assessment forward, flagged, rather than fabricate movement.
To compare origins at a common destination, we reconstruct the price chain layer by layer, each layer sourced:
EXW → + inland haulage → FOB → + ocean freight → CFR → + marine insurance → CIF → + destination haulage → DAP → + duty & fees → DDP
Origin differential and freight are estimates (labelled); insurance is a percentage of CFR; duty comes from the official schedule for the commodity’s HS code and destination. The freight leg works for any of ~1,600 ports via a distance model anchored on our lane rates.
We treat forecasting as an empirical discipline, not a black box. Three principles govern it:
Baseline first. Commodity spot prices are close to a random walk (weak-form efficiency), so the random-walk-with-drift is the benchmark every model must beat out-of-sample. We score models against it — MASE, RMSE, directional accuracy — and if a model does not beat the naive baseline, it is not published as signal.
A model that fits the data we actually have. Our series are monthly, sometimes gappy, and fuse public benchmarks with our own sparser observations — so our working model is a structural / state-space model estimated with the Kalman filter (local level or local linear trend, with seasonal terms). We choose it deliberately: it fuses multiple sources of different reliability as separate measurements, handles missing data natively, and nests the random walk, so it degrades gracefully rather than overfitting a short sample. Around it we test seasonal decomposition (STL) for harvest cycles, SARIMA / ARIMAX with exogenous drivers (freight, FX, inventories), error-correction (VECM) where a clean cointegrating partner exists (e.g. arabica–robusta), the GARCH family for volatility bands, and machine-learning challengers — each admitted only when it beats the baseline out-of-sample under walk-forward cross-validation (no look-ahead) and information criteria. Forecast combination(Bates–Granger) is the destination we grow into once several models each earn their place — not the starting point.
Forecast and decision are separate. We estimate and score a calibrated predictive distributionwith a symmetric proper scoring rule (CRPS / pinball) so models stay comparable; the buyer’s genuinely asymmetric cost (waiting and overpaying ≠ buying early into a fall) is applied at the decisionlayer, by reading the quantile of that distribution whose level equals the cost ratio — the newsvendor rule. That yields the asymmetric-optimal action without baking a fragile asymmetry into the estimator.
Uncertainty is first-class. Forecasts are probabilistic — intervals, not naked points — and we check interval calibration (coverage / PIT) out-of-sample.
Where we are now, honestly:we hold ~25 years of monthly public history for the core commodities — enough to backtest the baseline and state-space models properly — while our own transaction observations are still sparse and grow with every deal (the state-space model folds them in as they arrive). Richer structural and ML models switch on only when the backtest says they beat the random walk. We publish those backtests; we don’t dress up a model that hasn’t earned it.
CTX is indicative, not executable— the firm price is the desk’s quote. It is a trade reference, not a financial benchmark; were it ever used to settle financial contracts, that changes the regulatory picture and we would seek external and legal review first. The methodology is versioned; material changes are dated and change-logged. None of this is investment advice.
More commodities and corridors; denser own-transaction inputs; published rolling backtests against the baseline; and a documented data API. As the data deepens, so does the model — transparently.