Component handshakes, identifiers, preview gates, and service boundaries made the experience coherent in ways a polished screen alone could not.
Enterprise data science · o9 Digital Brain
DSML
Studio.
Extract. Transform. Load, and trust it.
A data scientist’s day is an ETL loop: find governed data, shape it, test a model, ship it, watch it run. I designed the workspace where that loop stays connected instead of scattering across six separate tools.
System design
- 01Discover
- 02Explore
- 03Build
- 04Validate
- 05Operate
- 06Observe
ReceivesProjectEnvironmentModelSelected dimensions + measures
Passes forwardProfiled dataset + access scope
- Problem
- Specialist tools.
Governed data.No shared thread between them. - My role
- Product Designer0→1 system design, six connected workspaces
- Worked with
- EngineeringArchitecture reviews and the implementation specification
- Status
- A design hypothesisNot measured in use yet. Validation is the next step.
Extract · 01 / Context
Why data science needed
a pipeline, not a pile of tools.
Data scientists worked in specialist tools while planners worked inside the governed enterprise model. Between them sat copied identifiers, brittle scripts, hidden assumptions, and delayed feedback.
How I framed it: how might we connect flexible data-science work to governed enterprise planning, without losing context at each handoff?
Model Explorer and Data Explorer pull from the live planning model or cloud files, never a silent copy.
EDA, notebooks, and plugins turn a raw frame into a versioned, reusable artifact.
Deployments and Observability put a result into production and keep it traceable when it fails.
A solid line: context carried from stage to stage.A reconstruction of the operating model, not a trace of observed sessions
Designed scope, not impact. These count what the design covers. None of them is a measured result.
Extract · 02 / Research approach
How the evidence became
product decisions.
I ran the same extract, transform, load discipline on the research itself: four inputs, one synthesis, three concrete decisions.
What this research is
Four inputs I could access: the implementation specification, the platform’s existing patterns, mature specialist tools, and architecture reviews with engineering.
What it isn’t
User interviews were not available for this phase. It established a design hypothesis, not a measured usability result.
Extract · 03 / Who it’s for
Four roles,
one shared model.
The same model passes through data science, engineering, planning, and operations. Each role needs different evidence before they can move the work forward.
Transform · 04 / Information architecture
Organised around
what the user means to do.
The navigation follows four goals: find data, build, operationalise, and monitor. The architecture underneath keeps governance and execution reusable.
Data foundation
Find + qualify
Model ExplorerData ExplorerMetadataBuild
Create + package
NotebooksPluginsParametersOperationalize
Compare + release
ExperimentsDeploymentsRegistryMonitor
Trace + recover
HardwareLogsRun historyGroups follow the product’s navigation; layers are named as in the implementation specification.
Transform · 05 / Follow the request
One model, from data to production.
Its context never retyped.
Six stages, three moves. Pick a stage to see the screen, the question it answers, the context it receives, and what it passes forward.
Six stages · one model’s context
Stage questions and handoffs are from the design’s journey; the global context bar (project, environment, model) is on every screen of the rebuild. Captures are from the Carbon rebuild, and all data is sample data.
Transform · 06 / Key decisions
Four decisions that made the workflow
safer and clearer.
Each adds a deliberate check before a consequential action, or removes repeated work where context can travel on its own.
Decision / 01
Separate authoring from execution
Decision 01 · ModeSystem state
Editing changes the reusable artifact; executing creates a run with scope and parameters. Mixing them would make save state, permissions, and failure recovery harder to understand.
The trade-off: one more explicit switch, far less ambiguity.
Mode clarity
Error prevention
Decision / 02
Make preview a gate, not a courtesy
Decision 02 · PreviewSystem state
PREVIEWExecute unlocks
The final query is composed from plugin config, scope, and parameters. Preview turns that hidden composition into evidence you can inspect before Execute becomes available.
The trade-off: a step before execution.
Safe exploration
Commitment threshold
Decision / 03
Preserve trace context automatically
Decision 03 · TraceIdentifiers · sample
PluginId, BatchJobId, and RequestId travel from execution into observability. The user should never have to copy an opaque identifier between tools.
The trade-off: it needs service-level handshake contracts.
Visibility
Recoverability
Decision / 04
Embed specialist tools without losing the shell
Decision 04 · ShellContext layers
JupyterLab, MLflow, and observability patterns stay recognisable, while the o9 shell keeps module, environment, role, and planning context stable.
The trade-off: integration complexity around theme and navigation.
Familiarity
Contextual integrity
Designing with data
Comparison only works when metrics answer a decision.
The experiment view avoids a “dashboard of everything.” It prioritises run identity, model family, accuracy, duration, owner, version, and status: the minimum set needed to register a candidate with confidence.
demand_forecast_w27 · first four runsRun records · sample
Load · 07 / Outcome & reflection
What changed, and what still
needs validation.
The six workspaces now form one understandable workflow. The next step is measuring whether teams complete cross-role tasks faster and recover from failures with less support.
What success should be measured by
Hypotheses · not yet measuredWhat I can say
The designs and the Carbon rebuild show six workspaces joined by carried context, a preview gate before execution, and request-level traceability into observability. The implementation specification defines the services underneath.
What I don’t claim
A launch, adoption, or any measured change in task time or failure recovery. User interviews weren’t available for this phase, and the run metrics shown are sample data.
Next, I’d measure time to a first governed dataset, failed-run recovery, confidence when registering a model, and what gets lost at cross-role handoffs.
Longitudinal task evidence.
Surface model lineage, feature contribution, and confidence where a planner consumes the writeback, not only inside DSML Studio.
Explainability at the planning moment.
Complex systems become usablewhen every handoff carries context.