ML for portfolios
Using machine learning to improve portfolios
- ML for portfolios
- Purpose
- Code and library base
- Use cases for machine learning in managing projects.
- Fitting ML into an organisation context
- Summary of start-up steps for application
- Applying machine learning at project or programme or portfolio levels
- More guidance on applying ML to portfolios
Purpose
Apply machine learning to understand how to improve the project portfolio
A worked example: project ratings in Orange
2019 exploration · explanation reviewed 1 October 2026
Could completed-project records help a portfolio manager decide where to look next? This World Bank example makes the process visible: prepare records, compare models, inspect the mistakes, and consider what a prediction might change in a monthly review.
Follow the top row from Extract to Sample data, then the lower branch into Test possible supervised models. The illustrated guide explains why a good-looking score and a useful management decision are different achievements.
Read the illustrated Orange walkthrough →
The original name was Project Success Prediction, but its target is IEG Bank Performance, a specific evaluation rating. This is a historical demonstration: the workbook is absent and the displayed accuracy has not been reproduced as an early forecast. The guide retains all 15 screenshots, the short film, demonstration script and saved workflow, with clear routes to each.
Film, script and workflow files · Back to the Library
Before work starts: the Highways delay notebook
Historical idea: August 2020 · synthetic reconstruction executed 4 October 2026
Could completed activities help a planner identify which unstarted work might finish late? The useful prior question is what information would actually be available when making that prediction. An actual finish date can explain the past while being unavailable for a forecast.
The public notebook now uses an entirely invented access-road programme. It replaces all original project rows, saved results and plots; every date, activity and number is synthetic. Lawrence Rowland’s 2020 Highways exploration supplies the question and learning sequence, not the data or results below.
- Separate what is known from what will happen. Generate 300 completed and 60 not-started activities from a fixed seed. Fit on the first 240 completed activities, check on the later 60, and leave the not-started outcomes unknown. The only predictors are stipulated pre-start fields: planned workdays, interfaces, access constraints and work type.
- Ask two different questions. A logistic classifier estimates whether finish variance reaches five days; a linear regressor estimates the variance itself. Here the convention is explicit: actual finish minus baseline finish, in days; positive means late. This convention belongs to the synthetic example, not the original source.
- Check against simple baselines. Compare the held-out results against the training majority class and mean variance. Keep activity identifiers attached when producing estimates for unstarted work. Training scores, held-out scores and future estimates are visibly different.
New synthetic plot, executed 4 October 2026. Both axes measure days of finish variance. The 60 held-out activities were generated by the same invented rule as the training set; this is a check of the demonstration, not evidence of forecasting ability on real projects.
What carries forward is the discipline of asking what we could know in time. The reconstruction excludes outcome, progress and identifier fields from model inputs, fits preparation on training rows only, and preserves row alignment. Those checks make the example easier to inspect. They do not prove the assumptions that a real forecast would need.
The historical exploration and this reconstruction are distinct. The original used DABL, lacked its source CSV, did not establish the finish-variance units, and left a separate test set and prediction-row checks unfinished. It has not been reproduced here. The new notebook uses small explicit NumPy models and a seeded fictional dataset; every displayed output was newly executed. Real use still needs permitted timestamped records, validation across projects or grouped activities, and evidence that the resulting decisions help.
Read or download the synthetic notebook →
The stable URL retains the historical notebook name. This is a tabular prediction example, separate from the neighbouring Neo4j graph-construction exercises. Reading guide · Information leakage · Back to the Library
A different question: monthly portfolio decisions
December 2020 working note · reading guide added 1 October 2026
What should we decide this month, and what should we learn before deciding again? This earlier note puts advancing, suspending and cancelling projects beside investigation, review and assurance. One set of choices changes the work; the other may change what we know about it.
Reading diagram made in 2026 from the original note. It illustrates the question; it does not calculate a policy.
The useful thread is state → decision → new information → next state and decision. Read S as the project’s position at a review, x as the decision, and W as information arriving afterwards. Does the timing of our reviews fit the information we need—and could finding out more change the next decision?
Read the monthly portfolio-review guide →
The two retained notebooks explore a much smaller idea: random choices between promote, maintain and cancel, with supplied rewards. Neither learns a policy, consumes new evidence or models assurance. The guide distinguishes their different attempts and known counting/transition defects; the original note and saved outputs remain available. This is an unfinished exploration, with useful questions rather than validated recommendations.
Original note and clarifications · Back to the Library
Code and library base
The guides and selected files above remain readable here. The broader GitHub repository contains other source material and may require access.
Use cases for machine learning in managing projects.
Below shows the phase of project delivery and Operations these use cases first appear.

Fitting ML into an organisation context
Twelve questions for you and your team
Retained from the earlier, undated ML questionnaire; moved here on 4 October 2026. Use the questions and suggested answers to discuss your starting point. There is nothing to submit, and this page does not collect or save your answers.
At what level do you mostly work when it comes to projects?
- At project, programme and portfolio level, also linking in with Strategy and Operations
- Portfolio level, down to individual projects
- Programme , project or work package level
- PMOs at various levels
- Other — note your own answer privately.
What is your understanding of ML?
- I don’t know how it differs from statistics, analysis and reporting
- I understand something of what it can be used for
- I’ve read a few articles and know the general principles
- I know the main types of machine learning
- I've done some courses/reading/prototypes but not applied it to projects
- Other — note your own answer privately.
Have you or the company used machine learning to run your projects better?
- Several times successfully
- One or twice, without much success
- No, but I’ve seen some specific examples I’d like to try
- No, but I’d like to know how to get started,
- Other — note your own answer privately.
Are there ML initiatives underway elsewhere?
- No
- Yes, in departments I work with
- Yes, in departments I don’t work with
- Not really, but there are some people with experience that can help
Are there ML frameworks, guidance or software within the company available to you?
Check all that apply
- Yes there is a framework I need to work within
- There is software or tools I should be using
- We have an arrangement with a supplier
- Other — note your own answer privately.
How many programmes and projects do you have?
- 1-10
- 10-100
- 100-1000
- Other — note your own answer privately.
For most projects, how many useful data fields do you have?
- Still focussing on tracking down which projects are live
- 1-2
- 3-5
- 5-10
- More than 10
Do you have consistent access to detailed project content and reports?
This question is focussing on how much content such as designs , reports , forms, assessments, correspondence is kept - this would be useful for natural language processing.
- Yes , mostly standardised documentation per project
- I can request access individually to a project area
- No, projects do their own thing
- No, projects don’t keep much textual content
- Other — note your own answer privately.
Do you have data across time?
Tick all which apply
- Yes, I have old project data going back years
- Yes, old projects are labelled successful, cancelled or unsuccessful
- Yes, we have detailed schedules per project
- Yes, we have time series data across projects tracking spend and status and delay
- No, it’s mostly just the current projects
- Other — note your own answer privately.
Do you have particular aspirations in the following areas?
- Project Success prediction
- Data exploration
- Enhanced search for similar projects / documents
- Project topic modelling
- Looking for strengths / weaknesses across schedule items
- Grouping projects within portfolio by similarity
What tends to be your preferred way of developing a new area?
- Commission a report or a study
- Bespoke workshops based upon my needs, spaced out so that I can apply the learning in-between
- Set things up together, learning as I go
- Commission a ML prototype or experiment
- Bring in an interim to your portfolio team to get this area started
- A few days coaching to get me started on the journey, and I can take it from there
- Other — note your own answer privately.
Do you have specific plans for how to use ML in improving your projects?
- I want to define or set up a programme of improvements
- I want to be able to intelligently procure ML services or products with a main supplier
- I want to understand in general what ML can do for projects before I make any plans
- I have identified a number of specific use cases, and have identified an ML approach
- I have identified some possible use cases, and want to know if they are reasonable
Summary of start-up steps for application
Once a use-case has been chosen, the following decisions can be made:
- Identify possible use cases
- Map to business area/lifecycle
- Narrow down to a generic ML method
- Select a straightforward ML model
- Choose a suitable model environment
- Choose a code library or algorithm to apply that model
- select data-set
More is said about each stage here
Applying machine learning at project or programme or portfolio levels
Machine learning provides different types of insight at different project levels. Some machine learning approaches make the most of the extra context provided by graph databases, specifically in terms of which relationships are meaningful.
-
PROJECT & PROGRAMME level: node, edge and property prediction for risks, dependencies, project sectoral properties and success.
-
PORTFOLIO level: mostly standard network algorithms (centrality, breadth first search etc). Includes the conversion between graph and tree structures to provide appropriate views for different stakeholders. The ML element is currently restricted to identifying common and anomalous graph motifs.
-
TASK AND SCHEDULE level: This is the least well developed, but a broad range of approaches to making the most of the Optimisation work of Professor Warren B. Powell, making the most of the inherent graph structure of resource-task-outcome paths. Eventually exploring Graph-Graph neural networks & Seq-Seq/ Transformer approaches as well as Monte Carlo Tree search
-
PMO and CENTRE OF EXCELLENCE. NLP applied to boost taxonomic and semantic approaches to curating body of project practice for the organisation.
More guidance on applying ML to portfolios
The guidance continues here

