Enterprise data science · o9 Digital Brain

DSML
Studio.

Extract. Transform. Load, and trust it.

A data scientist’s day is an ETL loop: find governed data, shape it, test a model, ship it, watch it run. I designed the workspace where that loop stays connected instead of scattering across six separate tools.

System design

DSML Studio / Data foundation / Exploratory data analysis
  1. 01Discover
  2. 02Explore
  3. 03Build
  4. 04Validate
  5. 05Operate
  6. 06Observe

ReceivesProjectEnvironmentModelSelected dimensions + measures

Passes forwardProfiled dataset + access scope

Problem
Specialist tools.
Governed data.No shared thread between them.
My role
Product Designer0→1 system design, six connected workspaces
Worked with
EngineeringArchitecture reviews and the implementation specification
Status
A design hypothesisNot measured in use yet. Validation is the next step.
Scroll to unpack

Extract · 01 / Context

Why data science needed
a pipeline, not a pile of tools.

Data scientists worked in specialist tools while planners worked inside the governed enterprise model. Between them sat copied identifiers, brittle scripts, hidden assumptions, and delayed feedback.

How I framed it: how might we connect flexible data-science work to governed enterprise planning, without losing context at each handoff?

ExtractFind governed data

Model Explorer and Data Explorer pull from the live planning model or cloud files, never a silent copy.

TransformClassify, model, package

EDA, notebooks, and plugins turn a raw frame into a versioned, reusable artifact.

LoadDeploy, trace, learn

Deployments and Observability put a result into production and keep it traceable when it fails.

Six handoffs, two operating modelsDSML Studio: the context travels with the work

    A solid line: context carried from stage to stage.A reconstruction of the operating model, not a trace of observed sessions

    6workspaces connectedone lifecycle
    12execution parametersoptional, typed, reusable
    3trace identifiersPlugin · Batch Job · Request
    5access levelsprivate to tenant-wide

    Designed scope, not impact. These count what the design covers. None of them is a measured result.

    Extract · 02 / Research approach

    How the evidence became
    product decisions.

    I ran the same extract, transform, load discipline on the research itself: four inputs, one synthesis, three concrete decisions.

    ExtractInputs I could access
    TransformHow I synthesised it
    Core taskMove a model from governed data to production
    Repeated riskContext disappears between tools, roles, and handoffs
    UX responseCarry context, preview commitments, trace every run
    LoadWhat changed in the design
    01Intent-based navigationFind, build, operationalise, monitor
    02Preview before actionShow data, parameters, impact first
    03Request-level traceabilityModel, version, environment, request ID

    What this research is

    Four inputs I could access: the implementation specification, the platform’s existing patterns, mature specialist tools, and architecture reviews with engineering.

    What it isn’t

    User interviews were not available for this phase. It established a design hypothesis, not a measured usability result.

    Extract · 03 / Who it’s for

    Four roles,
    one shared model.

    The same model passes through data science, engineering, planning, and operations. Each role needs different evidence before they can move the work forward.

    Transform · 04 / Information architecture

    Organised around
    what the user means to do.

    The navigation follows four goals: find data, build, operationalise, and monitor. The architecture underneath keeps governance and execution reusable.

    01

    Data foundation

    Find + qualify

    Model ExplorerData ExplorerMetadata
    02

    Build

    Create + package

    NotebooksPluginsParameters
    03

    Operationalize

    Compare + release

    ExperimentsDeploymentsRegistry
    04

    Monitor

    Trace + recover

    HardwareLogsRun history

    Groups follow the product’s navigation; layers are named as in the implementation specification.

    Transform · 05 / Follow the request

    One model, from data to production.
    Its context never retyped.

    Six stages, three moves. Pick a stage to see the screen, the question it answers, the context it receives, and what it passes forward.

    Six stages · one model’s context

    Prototype rebuildSample data
    DSML Studio

    Stage questions and handoffs are from the design’s journey; the global context bar (project, environment, model) is on every screen of the rebuild. Captures are from the Carbon rebuild, and all data is sample data.

    Transform · 06 / Key decisions

    Four decisions that made the workflow
    safer and clearer.

    Each adds a deliberate check before a consequential action, or removes repeated work where context can travel on its own.

    Decision / 01

    Separate authoring from execution

    Decision 01 · ModeSystem state

    EDITConfigCodeSave state
    EXECUTEScopeParametersRun state
    Why

    Editing changes the reusable artifact; executing creates a run with scope and parameters. Mixing them would make save state, permissions, and failure recovery harder to understand.

    The trade-off: one more explicit switch, far less ambiguity.

    Mode clarity

    Error prevention

    Decision / 02

    Make preview a gate, not a courtesy

    Decision 02 · PreviewSystem state

    Plugin configScopeParameters
    QUERY
    PREVIEW
    Execute unlocks
    Why

    The final query is composed from plugin config, scope, and parameters. Preview turns that hidden composition into evidence you can inspect before Execute becomes available.

    The trade-off: a step before execution.

    Safe exploration

    Commitment threshold

    Decision / 03

    Preserve trace context automatically

    Decision 03 · TraceIdentifiers · sample

    PluginId12346
    BatchJobId63ee…4a6
    RequestId82a1…0c7
    Why

    PluginId, BatchJobId, and RequestId travel from execution into observability. The user should never have to copy an opaque identifier between tools.

    The trade-off: it needs service-level handshake contracts.

    Visibility

    Recoverability

    Decision / 04

    Embed specialist tools without losing the shell

    Decision 04 · ShellContext layers

    o9 GLOBAL CONTEXT
    JupyterLabMLflowObservability
    Role · environment · model · scope
    Why

    JupyterLab, MLflow, and observability patterns stay recognisable, while the o9 shell keeps module, environment, role, and planning context stable.

    The trade-off: integration complexity around theme and navigation.

    Familiarity

    Contextual integrity

    Designing with data

    Comparison only works when metrics answer a decision.

    The experiment view avoids a “dashboard of everything.” It prioritises run identity, model family, accuracy, duration, owner, version, and status: the minimum set needed to register a candidate with confidence.

    Sample runsFrom the rebuild’s experiment table · not production performance

    demand_forecast_w27 · first four runsRun records · sample

    RMSE, lower is better → · R² ↑
    Candidate Other run
    Four sample runs, R squared against RMSE. demand_forecast_01, Random Forest: RMSE 18.4, R squared 0.940, the candidate. demand_forecast_02, XGBoost: 20.1, 0.928. demand_forecast_03, LightGBM: 21.8, 0.916. demand_forecast_04, Prophet: 23.5, 0.904. 0.940.920.90 18202224 _01 · Random Forest _02 · XGBoost _03 · LightGBM _04 · Prophet
    REGISTER CANDIDATEdemand_forecast_01 · Random ForestR² 0.940 · RMSE 18.4 · MAPE 6.8% · 2m 12s

    Load · 07 / Outcome & reflection

    What changed, and what still
    needs validation.

    The six workspaces now form one understandable workflow. The next step is measuring whether teams complete cross-role tasks faster and recover from failures with less support.

    What success should be measured by

    Hypotheses · not yet measured
    01DiscoverTime to a usable governed datasetNot measured
    02BuildContext switches per model iterationNot measured
    03ValidateTime to select and register a candidateNot measured
    04OperateMedian time from failure to causeNot measured

    What I can say

    The designs and the Carbon rebuild show six workspaces joined by carried context, a preview gate before execution, and request-level traceability into observability. The implementation specification defines the services underneath.

    What I don’t claim

    A launch, adoption, or any measured change in task time or failure recovery. User interviews weren’t available for this phase, and the run metrics shown are sample data.

    01

    Component handshakes, identifiers, preview gates, and service boundaries made the experience coherent in ways a polished screen alone could not.

    Design below the UI layer.

    02

    Next, I’d measure time to a first governed dataset, failed-run recovery, confidence when registering a model, and what gets lost at cross-role handoffs.

    Longitudinal task evidence.

    03

    Surface model lineage, feature contribution, and confidence where a planner consumes the writeback, not only inside DSML Studio.

    Explainability at the planning moment.

    Complex systems become usablewhen every handoff carries context.