Skip to content

Writing node code

A Function node is a Python file. That is all it is — no base class, no decorator, no framework import unless you want one.

def process(temperature, setpoint=21.0):
    """Ask for heat when the room is below the comfort point."""
    return {"heat": temperature < setpoint}

The rules

One function called process. If the file defines exactly one public function under another name, that one is used instead. Two, and the node refuses to load rather than guessing.

Arguments come from ports and settings, by name. temperature above is an input port; setpoint is a setting typed into the node's panel. Both arrive as keyword arguments, which is why a setting may not share a name with a port. See Where a node's values come from.

The return value is a dict keyed by output ports. Every value is checked against the port's declared type before it is published. A key that is not a declared port is an error, not a silent drop — nothing leaves a node except through a port it declared.

Nothing else is importable from the engine. Node code runs in a separate process, on a separate interpreter, with none of Fluksio's own modules on its path. What it can import is what the Modules screen installed.

A node is a pure function of its inputs. No context object, no global store, no handle to reach for. A running total or a debounce timer has a specific shape — see Keeping state in a flow.

Producing values over time

A node that produces values during its execution is a generator. Every yield is a dict keyed by output port, published the instant it happens:

def process(lr, steps):
    loss = 1.0
    for _ in range(steps):
        loss = train_one_step(lr)
        yield {"loss": loss}          # published now
    return {"final_loss": loss}

Whatever the generator returns at the end is the node's result — what downstream nodes read. If you never return, the last thing you yield is the result instead.

Mark the port so the flow says what it does:

{"name": "loss", "dtype": "float", "stream": true}

Two consequences. In a run, the whole series is kept as that run's metrics — this is why there is no log_metric() anywhere in the API. And the node's timeout starts measuring silence rather than duration: each emission resets the deadline, so a node yielding every few seconds can run for hours under a timeout of 300.

fluksio.emit

Where a yield cannot reach — the value comes from inside somebody else's callback, and they call you rather than the other way round:

import fluksio


def process():
    model.fit(callbacks=[LambdaCallback(
        on_epoch_end=lambda epoch, logs: fluksio.emit(loss=logs["loss"])
    )])
    return {"weights": ...}

Same ports, same type checking, same publication. Prefer yield where you can reach it; emit where you cannot.

Bytes: artifacts

Messages are JSON, which is what lets the same value pass through Redis, the work queue and the worker protocol unchanged. A checkpoint is not that.

import fluksio


def process(dataset):
    path = fluksio.load_artifact(dataset)     # → a local path to read
    ...
    return {
        "weights": fluksio.save_artifact("model.pt", media_type="application/octet-stream"),
        "score": 0.94,
    }

save_artifact takes bytes or a path, stores them by their SHA-256 digest, and returns a small reference — digest, size, media type, name — which is what an artifact-typed port carries.

Because the address is the content's hash, a sweep whose fifty configs share one preprocessed input stores it once, and a reference stays valid wherever the store is reachable from — including on another machine.

Printing

print works and is captured. The first 16 KB per call is kept and shown in the flow editor's log panel and on the run's per-node record; the rest is dropped, so a node printing in a loop cannot fill anything up.

Use it to debug. Do not use it to record results — a number worth keeping is an output port, not a line of text.

Errors

An exception fails that node's execution, not the flow. The message you see is one line from the frame in your code, not a stack through the engine — that is a deliberate choice about what is actionable.

The node keeps its last error visible after it recovers, so a failure that fired an alert at 03:00 still says what it was at 09:00. It can also be acknowledged from the canvas.

Timeouts

timeout on a node is how many seconds its code may run before it is stopped. The default is 30, and it covers the first call's imports, which can be much slower than the body — a node importing torch is not being slow, it is loading.

Above 60 seconds, a live flow may deliver the same work again while the node is still running. In a batch run, which never redelivers, it is an idle timeout instead: silence this long is a kill.

Running a node somewhere else

A node declares the label of the machine it needs:

{"id": "train", "device": "gpu", "device_policy": "require", "timeout": 7200}

require (the default) waits for a worker carrying that label; prefer runs locally when none is attached. A node bound to a device is compiled on that machine — a node importing torch is correct on the GPU box and a missing module on the engine, so checking it here would fail something that is fine.

See Remote workers.

Sharing code between flows

A node's source can be promoted to the shared library from its panel, and other flows can then use it by reference. One copy, one place to edit — and every flow using it runs the edit, which is the point and also the caution.

Shared sources live in _lib/ in the flow repository, so they are versioned with everything else.

Packages

Node code runs in a virtual environment of its own, on the installation's data volume — deliberately separate from the one Fluksio itself runs on.

Declare what you import in Modules, or over the API:

curl -X POST $FLUKSIO/modules/apply -H "Authorization: Bearer $TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"requirements": "numpy>=2\npandas\n"}'

It is a pip manifest installed with uv pip sync, versioned alongside your flows. An install takes effect immediately; nothing restarts.

A worked example

The repository ships a small supervised fit as a seedable demo — three nodes, a batch flow, streaming metrics, artifacts between stages, and a GPU-labelled node that falls back to the engine when no worker is attached. It is the shortest complete thing to read:

prepare ──dataset(artifact)──▶ train ──weights(artifact)──▶ evaluate
                                 └── loss (streaming float) ──▶ chart

make seed-demo builds it against a running stack.

See also