Cell 1 · markdown
Before work starts: a synthetic schedule-delay experiment
Historical idea: Lawrence Rowland, August 2020. Synthetic reconstruction: 4 October 2026.
The earlier Highways notebook asked whether completed activities could help identify delay in work not yet started. This public edition replaces its project records, saved outputs and plots with an entirely invented access-road example. Every row, date, score and plotted point below is synthetic. None describes an actual project, contract, person or place.
The original used DABL to explore classifiers and regressors. This reconstruction uses small explicit NumPy models so the data preparation, baseline and checks can be read and rerun without the missing source CSV or the historical environment. It preserves the learning sequence, not the original analysis or its results. The original file remains in restricted source history.
Run the cells in order with Python 3, NumPy, pandas and Pillow. No network service, database, local data file or credential is needed. The outputs in this edition were generated by executing the code on 4 October 2026. Re-execution may change formatting across library versions; the seeded data and model calculations are reproducible.
Cell 2 · code
import base64
import io
import math
import platform
import numpy as np
import pandas as pd
from PIL import Image, ImageDraw, ImageFont
SEED = 1932026
rng = np.random.default_rng(SEED)
print("Synthetic demonstration; seed", SEED)
print("Executed with Python", platform.python_version(), "NumPy", np.__version__, "pandas", pd.__version__)
Executed synthetic output · 4 October 2026
Synthetic demonstration; seed 1932026 Executed with Python 3.12.14 NumPy 2.3.5 pandas 2.2.3
Cell 3 · markdown
1. Make a fictional reporting snapshot
Imagine a small access-road programme with repeated work packages. A prediction is made before an activity starts. The only permitted predictors are planned workdays, the number of declared interfaces, the planned access-window constraint and the work type. These are stipulated to be known then; a real application must verify this against timestamped records.
The generator has 300 completed activities and 60 activities not started. Completion dates order the completed examples for a held-out later block. Finish variance is explicitly actual finish minus baseline finish, in days; positive is late. A delay flag means variance of at least five days. That convention is invented for this reconstruction; the original notebook did not establish its units and sign.
Cell 4 · code
n = 360
kinds = np.array(["Preparation", "Earthworks", "Drainage", "Paving"])
kind_index = rng.integers(0, len(kinds), n)
schedule = pd.DataFrame({
"ActivityID": [f"SYN-{i+1:03d}" for i in range(n)],
"WorkType": kinds[kind_index],
"PlannedWorkdays": rng.integers(2, 21, n),
"InterfaceCount": rng.integers(0, 6, n),
"LimitedAccess": rng.integers(0, 2, n),
"Status": ["Completed"] * 300 + ["Not Started"] * 60,
})
# Entirely invented outcome rule; no fitted coefficient comes from a real project.
synthetic_outcome = (
-2 + 0.18 * schedule["PlannedWorkdays"].to_numpy()
+ 1.1 * schedule["InterfaceCount"].to_numpy()
+ 3.2 * schedule["LimitedAccess"].to_numpy()
+ np.array([0.0, 1.0, 0.6, -0.3])[kind_index]
+ rng.normal(0, 2, n)
)
schedule["FinishVarianceDays"] = np.round(synthetic_outcome, 1)
# Do not retain even simulated future outcomes on the prediction-time table.
schedule.loc[schedule.Status == "Not Started", "FinishVarianceDays"] = np.nan
schedule["CompletionDate"] = pd.NaT
schedule.loc[:299, "CompletionDate"] = pd.date_range("2031-01-01", periods=300)
schedule["ReasonablyDelayed"] = pd.array(
[bool(v >= 5) if pd.notna(v) else pd.NA for v in schedule.FinishVarianceDays],
dtype="boolean",
)
del synthetic_outcome
print("Synthetic snapshot: 300 completed and 60 not-started activities.")
schedule.head(8)
Executed synthetic output · 4 October 2026
Synthetic snapshot: 300 completed and 60 not-started activities.
| ActivityID | WorkType | PlannedWorkdays | InterfaceCount | LimitedAccess | Status | FinishVarianceDays | CompletionDate | ReasonablyDelayed | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | SYN-001 | Paving | 2 | 2 | 0 | Completed | 0.3 | 2031-01-01 | False |
| 1 | SYN-002 | Paving | 3 | 0 | 0 | Completed | -0.1 | 2031-01-02 | False |
| 2 | SYN-003 | Earthworks | 10 | 3 | 1 | Completed | 9.3 | 2031-01-03 | True |
| 3 | SYN-004 | Earthworks | 2 | 5 | 1 | Completed | 6.3 | 2031-01-04 | True |
| 4 | SYN-005 | Paving | 9 | 2 | 0 | Completed | 1.9 | 2031-01-05 | False |
| 5 | SYN-006 | Earthworks | 8 | 1 | 0 | Completed | 3.1 | 2031-01-06 | False |
| 6 | SYN-007 | Earthworks | 13 | 2 | 0 | Completed | 4.0 | 2031-01-07 | False |
| 7 | SYN-008 | Drainage | 14 | 4 | 1 | Completed | 10.8 | 2031-01-08 | True |
Cell 5 · markdown
2. Separate fitting, later checks and prediction-time rows
Use the first 240 completed rows for fitting and the next 60 for a later held-out check. The last 60 have no outcome available and receive predictions only. The split is deliberately visible and preserves activity identifiers.
The invented generator is stationary: the held-out block is later in the fictional timeline but was drawn from the same rule. Passing it demonstrates the calculation under that rule, not real-world generalisation. Real schedules also need controls for repeated activity snapshots, related packages, projects and changing reporting practice.
Cell 6 · code
completed = schedule.loc[schedule.Status == "Completed"].sort_values("CompletionDate").copy()
training = completed.iloc[:240].copy()
held_out = completed.iloc[240:].copy()
not_started = schedule.loc[schedule.Status == "Not Started"].copy()
assert set(training.ActivityID).isdisjoint(held_out.ActivityID)
assert set(completed.ActivityID).isdisjoint(not_started.ActivityID)
assert not_started.FinishVarianceDays.isna().all()
assert not_started.ReasonablyDelayed.isna().all()
pd.DataFrame({"Synthetic block": ["Fit", "Held-out completed", "Not started"],
"Rows": [len(training), len(held_out), len(not_started)],
"Has observed outcome in this example": [True, True, False]})
Executed synthetic output · 4 October 2026
| Synthetic block | Rows | Has observed outcome in this example | |
|---|---|---|---|
| 0 | Fit | 240 | True |
| 1 | Held-out completed | 60 | True |
| 2 | Not started | 60 | False |
Cell 7 · markdown
3. Inspect the completed data without leaking its outcome
A completed-only summary helps us ask whether the assumed relationships are visible. It is not evidence that work type causes delay. Actual finish, variance, the derived delay flag, status and completion date are never model inputs. Nor is an activity identifier a meaningful predictor.
A real reporting export may call a field “planned” even though it was revised after the prediction date. Naming the inputs does not prove their availability; that needs a record-level audit.
Cell 8 · code
training.groupby("WorkType", sort=True).agg(
SyntheticTrainingRows=("ActivityID", "size"),
MeanVarianceDays=("FinishVarianceDays", "mean"),
ShareAtLeastFiveDaysLate=("ReasonablyDelayed", "mean"),
).round(3)
Executed synthetic output · 4 October 2026
| SyntheticTrainingRows | MeanVarianceDays | ShareAtLeastFiveDaysLate | |
|---|---|---|---|
| WorkType | |||
| Drainage | 73 | 4.799 | 0.479 |
| Earthworks | 68 | 5.590 | 0.588 |
| Paving | 46 | 4.141 | 0.435 |
| Preparation | 53 | 4.785 | 0.472 |
Cell 9 · markdown
4. Prepare features using training rows only
Work type is represented by indicator columns, with Preparation as the reference. Standardisation is fitted only to the training block and reused unchanged. This prevents information from the held-out block entering preparation. Feature selection is a stated judgement, not a claim that these are sufficient causes.
Cell 10 · code
FEATURES = ["PlannedWorkdays", "InterfaceCount", "LimitedAccess"]
def raw_features(frame):
numeric = frame[FEATURES].to_numpy(dtype=float)
indicators = np.column_stack([(frame.WorkType == k).to_numpy(dtype=float)
for k in kinds[1:]])
return np.column_stack([numeric, indicators])
train_raw = raw_features(training)
centre = train_raw.mean(axis=0)
scale = train_raw.std(axis=0)
scale[scale == 0] = 1.0
def features(frame):
z = (raw_features(frame) - centre) / scale
return np.column_stack([np.ones(len(frame)), z])
X_train, X_test, X_future = map(features, [training, held_out, not_started])
y_train = training.FinishVarianceDays.to_numpy(dtype=float)
y_test = held_out.FinishVarianceDays.to_numpy(dtype=float)
c_train = training.ReasonablyDelayed.to_numpy(dtype=float)
c_test = held_out.ReasonablyDelayed.to_numpy(dtype=bool)
assert X_train.shape == (240, 7)
print("Synthetic fit matrix:", X_train.shape, "including intercept; no outcome fields.")
Executed synthetic output · 4 October 2026
Synthetic fit matrix: (240, 7) including intercept; no outcome fields.
Cell 11 · markdown
5. Fit a classifier and a regressor
The classifier is logistic regression, fitted by a fixed gradient-descent recipe. The regressor is a least-squares linear model. They answer different questions: whether the stipulated five-day threshold will be crossed, and how many days of variance to estimate.
The classifier’s number is a model score in this synthetic setting. It has not been calibrated as a real project’s risk probability. A threshold of 0.5 is chosen for illustration; a manager would need to weigh missed delays against needless intervention.
Cell 12 · code
def sigmoid(z):
return 1 / (1 + np.exp(-np.clip(z, -35, 35)))
classifier = np.zeros(X_train.shape[1])
for _ in range(2000):
gradient = X_train.T @ (sigmoid(X_train @ classifier) - c_train) / len(c_train)
classifier -= 0.08 * gradient
regressor, *_ = np.linalg.lstsq(X_train, y_train, rcond=None)
train_score = sigmoid(X_train @ classifier)
test_score = sigmoid(X_test @ classifier)
train_variance = X_train @ regressor
test_variance = X_test @ regressor
assert np.isfinite(classifier).all() and np.isfinite(regressor).all()
print("Two models fitted to 240 synthetic completed activities.")
Executed synthetic output · 4 October 2026
Two models fitted to 240 synthetic completed activities.
Cell 13 · markdown
6. Compare against simple baselines on the held-out block
A majority-class baseline always makes the training block’s most common decision. A mean-variance baseline always predicts its average variance. Report accuracy and mean absolute error separately; one cannot substitute for the other. Training scores are shown beside, and distinguished from, held-out scores.
Cell 14 · code
majority = bool(c_train.mean() >= 0.5)
mean_variance = float(y_train.mean())
metrics = pd.DataFrame([
{"Synthetic check": "Training (not validation)",
"Classifier accuracy": np.mean((train_score >= 0.5) == c_train.astype(bool)),
"Majority accuracy": np.mean(majority == c_train.astype(bool)),
"Regression MAE (days)": np.mean(np.abs(train_variance - y_train)),
"Mean baseline MAE (days)": np.mean(np.abs(mean_variance - y_train))},
{"Synthetic check": "Held-out later completed block",
"Classifier accuracy": np.mean((test_score >= 0.5) == c_test),
"Majority accuracy": np.mean(majority == c_test),
"Regression MAE (days)": np.mean(np.abs(test_variance - y_test)),
"Mean baseline MAE (days)": np.mean(np.abs(mean_variance - y_test))},
]).set_index("Synthetic check")
metrics.round(3)
Executed synthetic output · 4 October 2026
| Classifier accuracy | Majority accuracy | Regression MAE (days) | Mean baseline MAE (days) | |
|---|---|---|---|---|
| Synthetic check | ||||
| Training (not validation) | 0.808 | 0.500 | 1.482 | 2.660 |
| Held-out later completed block | 0.750 | 0.467 | 1.760 | 2.875 |
Cell 15 · code
predicted_flag = test_score >= 0.5
pd.DataFrame({
"Synthetic held-out outcome": ["Below five days", "At least five days"],
"Predicted below five days": [int(np.sum(~c_test & ~predicted_flag)), int(np.sum(c_test & ~predicted_flag))],
"Predicted at least five days": [int(np.sum(~c_test & predicted_flag)), int(np.sum(c_test & predicted_flag))],
})
Executed synthetic output · 4 October 2026
| Synthetic held-out outcome | Predicted below five days | Predicted at least five days | |
|---|---|---|---|
| 0 | Below five days | 26 | 6 |
| 1 | At least five days | 9 | 19 |
Cell 16 · markdown
7. See what the held-out regression does
Each point is one invented held-out activity. The diagonal means estimate equals generated outcome. Vertical distance from that line is error. Unlike the earlier notebook’s training comparison and ratio plot, this uses withheld outcomes and a common scale; dividing by an outcome near zero would be unstable.
Cell 17 · code
def regression_plot(actual, predicted):
image = Image.new("RGB", (760, 490), "#f8f5ed")
draw = ImageDraw.Draw(image)
font = ImageFont.load_default(size=18)
small = ImageFont.load_default(size=15)
draw.text((32, 18), "SYNTHETIC · held-out schedule variance", fill="#173f3a", font=font)
draw.text((32, 46), "60 invented activities; no real project records", fill="#51595a", font=small)
left, top, right, bottom = 84, 92, 700, 397
lo = 5 * math.floor(min(actual.min(), predicted.min()) / 5)
hi = 5 * math.ceil(max(actual.max(), predicted.max()) / 5)
sx = lambda x: left + (x-lo)/(hi-lo)*(right-left)
sy = lambda y: bottom - (y-lo)/(hi-lo)*(bottom-top)
for tick in range(lo, hi+1, 5):
draw.line((sx(tick), top, sx(tick), bottom), fill="#dddcd3", width=1)
draw.line((left, sy(tick), right, sy(tick)), fill="#dddcd3", width=1)
draw.text((sx(tick)-7, bottom+10), str(tick), fill="#343c3b", font=small)
draw.text((left-35, sy(tick)-8), str(tick), fill="#343c3b", font=small)
draw.line((left, bottom, right, top), fill="#ba7942", width=2)
draw.rectangle((left, top, right, bottom), outline="#173f3a", width=2)
for a, p in zip(actual, predicted):
x, y = sx(a), sy(p)
draw.ellipse((x-4, y-4, x+4, y+4), fill="#247a70", outline="#173f3a")
draw.text((left, top-23), "Predicted finish variance (days)", fill="#173f3a", font=small)
draw.text((180, 443), "Generated finish variance (days; positive = late)", fill="#173f3a", font=small)
return image
held_out_plot = regression_plot(y_test, test_variance)
held_out_plot
Executed synthetic output · 4 October 2026
Cell 18 · markdown
8. Predict for not-started activities without pretending to know their outcomes
Attach both estimates using the preserved activity identifier. These rows deliberately have no actual result, so there is no future-performance score. Ranking a work package for review is a decision aid; it does not show that an intervention would prevent a delay. The model does not infer a causal policy or build a resource-feasible schedule.
Cell 19 · code
future = not_started.set_index("ActivityID")[
["WorkType", "PlannedWorkdays", "InterfaceCount", "LimitedAccess"]
].copy()
future["Synthetic delay score"] = sigmoid(X_future @ classifier)
future["Predicted variance days"] = X_future @ regressor
assert future.index.is_unique
assert list(future.index) == list(not_started.ActivityID)
future.round(3).head(8)
Executed synthetic output · 4 October 2026
| WorkType | PlannedWorkdays | InterfaceCount | LimitedAccess | Synthetic delay score | Predicted variance days | |
|---|---|---|---|---|---|---|
| ActivityID | ||||||
| SYN-301 | Drainage | 20 | 1 | 1 | 0.726 | 6.449 |
| SYN-302 | Drainage | 7 | 3 | 1 | 0.738 | 6.089 |
| SYN-303 | Earthworks | 7 | 1 | 0 | 0.059 | 1.426 |
| SYN-304 | Preparation | 3 | 3 | 0 | 0.115 | 2.119 |
| SYN-305 | Drainage | 3 | 1 | 0 | 0.024 | 0.192 |
| SYN-306 | Preparation | 8 | 1 | 0 | 0.037 | 0.864 |
| SYN-307 | Preparation | 11 | 0 | 0 | 0.022 | 0.338 |
| SYN-308 | Drainage | 12 | 5 | 0 | 0.819 | 6.540 |
Cell 20 · markdown
What carries forward from the 2020 exploration
- Ask what the prediction is for and which information exists at that moment.
- Distinguish classification, numeric prediction, a causal claim and a management decision.
- Keep activity identifiers attached; a plausible number against the wrong row is still wrong.
- Fit preparation on training data, then check on a separate relevant block and against a simple baseline.
- State units, reporting conventions, missing outcomes and the limits of any plot or score.
This reconstruction fixes several demonstration gaps while deliberately making no claim about historical model quality or real-project forecasting. The synthetic rule is learnable because it was designed that way. The original DABL model-selection experiment is part of the historical idea, not an algorithm reproduced here. Real deployment would require permitted timestamped data, grouped or project-level validation, distribution-change checks and an evaluation of whether the resulting decisions help.