ML for portfolios

Using machine learning to improve portfolios

  1. ML for portfolios
  2. Purpose
    1. A worked example: project ratings in Orange
    2. Before work starts: the Highways delay notebook
    3. A different question: monthly portfolio decisions
  3. Code and library base
  4. Use cases for machine learning in managing projects.
  5. Fitting ML into an organisation context
  6. Summary of start-up steps for application
  7. Applying machine learning at project or programme or portfolio levels
  8. More guidance on applying ML to portfolios

Purpose

Apply machine learning to understand how to improve the project portfolio

A worked example: project ratings in Orange

2019 exploration · explanation reviewed 1 October 2026

Could completed-project records help a portfolio manager decide where to look next? This World Bank example makes the process visible: prepare records, compare models, inspect the mistakes, and consider what a prediction might change in a monthly review.

Original Orange workflow: project data preparation branches into model comparison, visual exploration and an intended prediction route.

Follow the top row from Extract to Sample data, then the lower branch into Test possible supervised models. The illustrated guide explains why a good-looking score and a useful management decision are different achievements.

Read the illustrated Orange walkthrough →

The original name was Project Success Prediction, but its target is IEG Bank Performance, a specific evaluation rating. This is a historical demonstration: the workbook is absent and the displayed accuracy has not been reproduced as an early forecast. The guide retains all 15 screenshots, the short film, demonstration script and saved workflow, with clear routes to each.

Film, script and workflow files · Back to the Library

Before work starts: the Highways delay notebook

Historical idea: August 2020 · synthetic reconstruction executed 4 October 2026

Could completed activities help a planner identify which unstarted work might finish late? The useful prior question is what information would actually be available when making that prediction. An actual finish date can explain the past while being unavailable for a forecast.

The public notebook now uses an entirely invented access-road programme. It replaces all original project rows, saved results and plots; every date, activity and number is synthetic. Lawrence Rowland’s 2020 Highways exploration supplies the question and learning sequence, not the data or results below.

  1. Separate what is known from what will happen. Generate 300 completed and 60 not-started activities from a fixed seed. Fit on the first 240 completed activities, check on the later 60, and leave the not-started outcomes unknown. The only predictors are stipulated pre-start fields: planned workdays, interfaces, access constraints and work type.
  2. Ask two different questions. A logistic classifier estimates whether finish variance reaches five days; a linear regressor estimates the variance itself. Here the convention is explicit: actual finish minus baseline finish, in days; positive means late. This convention belongs to the synthetic example, not the original source.
  3. Check against simple baselines. Compare the held-out results against the training majority class and mean variance. Keep activity identifiers attached when producing estimates for unstarted work. Training scores, held-out scores and future estimates are visibly different.

Synthetic held-out scatter plot comparing generated and predicted finish variance for 60 invented activities; the diagonal means equality.

New synthetic plot, executed 4 October 2026. Both axes measure days of finish variance. The 60 held-out activities were generated by the same invented rule as the training set; this is a check of the demonstration, not evidence of forecasting ability on real projects.

What carries forward is the discipline of asking what we could know in time. The reconstruction excludes outcome, progress and identifier fields from model inputs, fits preparation on training rows only, and preserves row alignment. Those checks make the example easier to inspect. They do not prove the assumptions that a real forecast would need.

The historical exploration and this reconstruction are distinct. The original used DABL, lacked its source CSV, did not establish the finish-variance units, and left a separate test set and prediction-row checks unfinished. It has not been reproduced here. The new notebook uses small explicit NumPy models and a seeded fictional dataset; every displayed output was newly executed. Real use still needs permitted timestamped records, validation across projects or grouped activities, and evidence that the resulting decisions help.

Read or download the synthetic notebook →

The stable URL retains the historical notebook name. This is a tabular prediction example, separate from the neighbouring Neo4j graph-construction exercises. Reading guide · Information leakage · Back to the Library

A different question: monthly portfolio decisions

December 2020 working note · reading guide added 1 October 2026

What should we decide this month, and what should we learn before deciding again? This earlier note puts advancing, suspending and cancelling projects beside investigation, review and assurance. One set of choices changes the work; the other may change what we know about it.

This month's state and decision lead through progress and new information into next month's updated state and decision.

Reading diagram made in 2026 from the original note. It illustrates the question; it does not calculate a policy.

The useful thread is state → decision → new information → next state and decision. Read S as the project’s position at a review, x as the decision, and W as information arriving afterwards. Does the timing of our reviews fit the information we need—and could finding out more change the next decision?

Read the monthly portfolio-review guide →

The two retained notebooks explore a much smaller idea: random choices between promote, maintain and cancel, with supplied rewards. Neither learns a policy, consumes new evidence or models assurance. The guide distinguishes their different attempts and known counting/transition defects; the original note and saved outputs remain available. This is an unfinished exploration, with useful questions rather than validated recommendations.

Original note and clarifications · Back to the Library

Code and library base

The guides and selected files above remain readable here. The broader GitHub repository contains other source material and may require access.

Use cases for machine learning in managing projects.

Below shows the phase of project delivery and Operations these use cases first appear.

Fitting ML into an organisation context

Twelve questions for you and your team

Retained from the earlier, undated ML questionnaire; moved here on 4 October 2026. Use the questions and suggested answers to discuss your starting point. There is nothing to submit, and this page does not collect or save your answers.

  1. At what level do you mostly work when it comes to projects?

    • At project, programme and portfolio level, also linking in with Strategy and Operations
    • Portfolio level, down to individual projects
    • Programme , project or work package level
    • PMOs at various levels
    • Other — note your own answer privately.
  2. What is your understanding of ML?

    • I don’t know how it differs from statistics, analysis and reporting
    • I understand something of what it can be used for
    • I’ve read a few articles and know the general principles
    • I know the main types of machine learning
    • I've done some courses/reading/prototypes but not applied it to projects
    • Other — note your own answer privately.
  3. Have you or the company used machine learning to run your projects better?

    • Several times successfully
    • One or twice, without much success
    • No, but I’ve seen some specific examples I’d like to try
    • No, but I’d like to know how to get started,
    • Other — note your own answer privately.
  4. Are there ML initiatives underway elsewhere?

    • No
    • Yes, in departments I work with
    • Yes, in departments I don’t work with
    • Not really, but there are some people with experience that can help
  5. Are there ML frameworks, guidance or software within the company available to you?

    Check all that apply

    • Yes there is a framework I need to work within
    • There is software or tools I should be using
    • We have an arrangement with a supplier
    • Other — note your own answer privately.
  6. How many programmes and projects do you have?

    • 1-10
    • 10-100
    • 100-1000
    • Other — note your own answer privately.
  7. For most projects, how many useful data fields do you have?

    • Still focussing on tracking down which projects are live
    • 1-2
    • 3-5
    • 5-10
    • More than 10
  8. Do you have consistent access to detailed project content and reports?

    This question is focussing on how much content such as designs , reports , forms, assessments, correspondence is kept - this would be useful for natural language processing.

    • Yes , mostly standardised documentation per project
    • I can request access individually to a project area
    • No, projects do their own thing
    • No, projects don’t keep much textual content
    • Other — note your own answer privately.
  9. Do you have data across time?

    Tick all which apply

    • Yes, I have old project data going back years
    • Yes, old projects are labelled successful, cancelled or unsuccessful
    • Yes, we have detailed schedules per project
    • Yes, we have time series data across projects tracking spend and status and delay
    • No, it’s mostly just the current projects
    • Other — note your own answer privately.
  10. Do you have particular aspirations in the following areas?

    • Project Success prediction
    • Data exploration
    • Enhanced search for similar projects / documents
    • Project topic modelling
    • Looking for strengths / weaknesses across schedule items
    • Grouping projects within portfolio by similarity
  11. What tends to be your preferred way of developing a new area?

    • Commission a report or a study
    • Bespoke workshops based upon my needs, spaced out so that I can apply the learning in-between
    • Set things up together, learning as I go
    • Commission a ML prototype or experiment
    • Bring in an interim to your portfolio team to get this area started
    • A few days coaching to get me started on the journey, and I can take it from there
    • Other — note your own answer privately.
  12. Do you have specific plans for how to use ML in improving your projects?

    • I want to define or set up a programme of improvements
    • I want to be able to intelligently procure ML services or products with a main supplier
    • I want to understand in general what ML can do for projects before I make any plans
    • I have identified a number of specific use cases, and have identified an ML approach
    • I have identified some possible use cases, and want to know if they are reasonable

Summary of start-up steps for application

Once a use-case has been chosen, the following decisions can be made:

More is said about each stage here

Applying machine learning at project or programme or portfolio levels

Machine learning provides different types of insight at different project levels. Some machine learning approaches make the most of the extra context provided by graph databases, specifically in terms of which relationships are meaningful.

More guidance on applying ML to portfolios

The guidance continues here