What’s New in FLARE v2.9.0
Highlights:
Agent Skills — agent-assisted federated development
Collaboration API — a Python-first API for research workflows
Slurm job launcher — a new HPC execution target alongside process, Docker, and Kubernetes
Large-model training — a hardened model-transfer streaming transport and FedAvg validated to 72 billion parameters
Security hardening — authenticated CellNet messages, internal mTLS by default, and hardened admin and job-signing paths
Kubernetes/OpenShift deployment and framework/recipe additions also shipped this release; see Also in This Release below.
Agent Skills
FLARE Agent Skills fall into two categories:
Conversion skills generate a reviewable federated job from an existing project or dataset:
Training conversion — PyTorch, PyTorch Lightning, and Hugging Face Trainer are currently supported — identifies the owning framework, preserves the training and evaluation semantics, generates the supported Client API or recipe integration, validates the generated artifact, and reports evidence.
Federated statistics (tabular and image) generates a
FedStatsRecipejob directly from the dataset and feature names, with no user statistics code required.
Auto-FL optimization — an agent-directed campaign that tunes an existing job within its declared training budget:
NVFLARE owns the deterministic campaign import, execution, policy boundaries, and provenance.
The coding agent proposes hypothesis-driven candidates, constrained to the job’s fixed training budget and mutation-schema bounds.
Auto-FL’s initial importer supports statically recognizable NVFLARE Recipe and
*Jobpatterns.
Bundled skills are validated by pre-merge security scans, including prompt-injection and untrusted-input eval coverage, and include explicit safeguards for site-local data and preprocessing. Install and invoke them through a coding agent as described in Agent Skills (see NVFlare Auto-FL Agent Skill for the Auto-FL workflow); start with the runnable Agent Skills examples to try the conversion and federated-statistics workflows.
Agent Skills are developer tooling, not a runtime FL API — review a generated job before running it.
Collaboration API
Technical Preview
The Collaboration API is a technical preview designed for researchers to run quick experiments. It can run and deploy on a real multi-machine setup, but is not recommended for production at this rollout.
The Collaboration API provides a Python-first way to express custom federated
algorithms: decorate the functions that a server or client publishes, write
the coordination logic in ordinary Python, and use CollabRecipe to
package, export, simulate, or submit the result. This suits research
workflows that don’t fit a standard controller pattern. Every Collab call is
now authorized against the caller’s authenticated CellNet origin before
dispatch, rejecting a caller, method, or target that doesn’t match the call
envelope.
A simplified sketch of the API’s shape, not a literal excerpt — the client
publishes an ordinary method, and the server calls it on every client as if
it were local, with no Shareable, DXO, or FLModel transport
objects. See the runnable hello-collab example linked below for the
complete version:
from nvflare.collab import CollabRecipe, collab
from nvflare.recipe import SimEnv
class Trainer:
@collab.publish # publishes this method to clients under the name "train"
def train(self, weights=None):
... # local training
return updated_weights, loss
class FedAvg:
@collab.main
def run(self):
weights = None
for _ in range(num_rounds):
# "train" here calls the published train() method on every client
results = collab.clients.train(weights)
weights = average(results)
return weights
recipe = CollabRecipe(job_name="hello_fedavg", server=FedAvg(), client=Trainer())
recipe.execute(SimEnv(num_clients=2))
Note
This sketch illustrates the key code flow only. For the complete,
runnable code, see the hello-collab example below.
New examples:
Hello Collab — a minimal FedAvg workflow
pt_async_cifar10 — FedBuff-style buffered asynchronous aggregation at scale (up to 1,000 logical clients)
advanced Collab examples — split learning, swarm learning, and in-time aggregation
Slurm, Kubernetes, and Docker Job Launcher
FLARE 2.9.0 adds a new Slurm job launcher for HPC environments, joining the existing process, Docker, and Kubernetes launchers, and hardens the Docker and Kubernetes launchers with mTLS-by-default internal links and restricted job-controlled options (see Compatibility and Migration Notes). A long-lived NVFLARE parent submits each client or server job process as a Slurm batch job; Slurm selects resources while FLARE manages the federated job lifecycle. The Slurm launcher supports:
Apptainer, Pyxis/Enroot, and bare-Python execution backends
GPU-aware worker setup and multi-node applications
a shared-file worker channel for clusters where compute nodes cannot open a direct connection to the parent
Follow the Slurm Job Launcher deployment guide for prerequisites, backend setup, site configuration, and validation steps.
Large-Model Training
FLARE 2.9.0 strengthens the streaming transport used for large model transfers, across three areas:
Reliable Streaming — a transfer survives interruptions instead of failing outright:
Unacknowledged chunks retry within bounded retry budgets, and receiver-confirmed completion holds a payload until the receiver has actually consumed it.
A progress-aware liveness policy keeps a task download or result upload alive as long as bytes keep advancing, instead of failing or resending on a fixed wall-clock timeout; a transfer that truly stalls still fails after the configured idle limit. This now extends to Swarm Learning, where the aggregation client’s result-upload progress is tracked by its exact FQCN in relay and hierarchical topologies instead of falling back to single-receiver progress tracking.
External-trainer task materialization no longer trips a heartbeat expiry while a
TASK_READYexchange is pending (an optional task-wait timeout still bounds it).
Throughput and flow control — sender and receiver stay in sync under load:
The sender now tells the receiver its effective chunk and window size on every frame, and its ACK interval (plus retry wait/timeout for reliable streams) on the first frame of each stream, so mismatched endpoint settings can’t stall flow control.
Receiver reassembly capacity tracks the negotiated stream window instead of a fixed chunk count, so scheduler-induced chunk reordering doesn’t abort healthy transfers under load.
Pipelined tensor downloads, prefetching, and
TCP_NODELAYimprove throughput; oversized blobs fail before transmission with an actionable error, and a failed streamed result send now retries instead of silently dropping the result.Active administrative result downloads refresh their bound HCI session as bytes advance, so a healthy long download doesn’t expire as idle.
Memory — peak aggregator RSS stays flatter as models and client counts grow (the FL server for FedAvg/Scaffold/FedOpt; the client-elected aggregator site for Swarm):
Tensor disk offload during aggregation, previously FedAvg-only, now also covers Scaffold, FedOpt, and Swarm.
Pass-through tensor broadcasts release their source transaction as soon as downstream consumers finish, instead of retaining a model-sized object per aggregation round.
The default maximum streamed blob size is 4 GiB, and remains configurable.
With suitable infrastructure and configuration, FedAvg has been validated for federated LLM training at scales up to 72 billion parameters. See Large Models for deployment sizing and large-model operational guidance.
Training time and server memory across model sizes (1.7B-72B)
Elapsed time, 1.7B-72B (measured, 1 FL round) |
FedAvg server peak memory, 1.7B-72B (measured, 1 FL round) |
Each configuration in both charts ran a single FL round with two simultaneous clients and four local optimizer steps per client; the measurements characterize per-round transfer time and server peak memory, not a full convergence training run.
With pass-through download on, Swarm Learning’s tensor disk offload lowers the fixed aggregator’s peak memory, and the savings widen with model size; non-aggregator sites stay approximately flat at every size, since disk offload targets contribution handling at the aggregator rather than the per-site learner footprint (external-process, fixed aggregator, 4 clients, 30 rounds):
Swarm Learning aggregator peak memory reduction (with disk offload)
Aggregator (site-1) peak container memory, disk offload OFF vs. ON:
Model |
OFF peak |
ON peak |
Reduction |
|---|---|---|---|
5 GB synthetic |
48.49 GiB |
37.08 GiB |
23.5% |
30 GB (Qwen2.5-14B) |
236.80 GiB |
150.80 GiB |
36.3% |
60 GB synthetic |
452.20 GiB |
287.40 GiB |
36.4% |
Non-aggregator sites (site-2/3/4) moved by -1.9% to +3.0% across all three model sizes – run-to-run peak variation, not a disk-offload effect.
Security Hardening
FLARE 2.9.0 hardens the internal transport and admin access, on top of moving job-process bootstrap credentials off the command line (see Compatibility and Migration Notes below for that migration’s requirements):
CellNet message authentication. Cell payload encryption moves from unauthenticated AES-CBC to signed AES-256-GCM envelopes, and the sender signature on every message — including cached-key paths — is now verified before it’s trusted, closing a ciphertext bit-flipping exposure. This is a wire-format change: a 2.9 peer rejects the legacy unversioned ciphertext, so encrypted CellNet participants must upgrade to 2.9 together.
Internal mTLS by default. Internal CellNet TCP links between a parent and its job processes now default to mutual TLS across Docker, Slurm, Kubernetes, and Network Attach deployments, each with an explicit clear-transport opt-out for sites that intentionally run without it.
Certless admin session and listener hardening. An admin session token that can’t be signature-verified now fails closed instead of falling back; admin, TCP, and SimEnv listeners bind to explicit loopback or configured hosts instead of a broad wildcard default.
Cross-client authentication now routes through the server. A message between different client families is authenticated through the server trust boundary even when a direct or cached peer endpoint would otherwise be used.
``require_signed_jobs`` is now also enforced client-side. The policy itself shipped server-side in 2.8; 2.9 adds the same enforcement at the receiving client, rejecting unsigned job deployment bytes there too. Exact-byte signature verification is preserved even when unsigned jobs are otherwise allowed.
CLI, diagnostics, and Recipe secret handling. CLI and runtime diagnostics redact sensitive values more consistently, and Recipe APIs add safeguards for declaring and handling secrets; see Keeping Secrets Out Of Recipe Parameters.
Also in This Release
Kubernetes and OpenShift deployment — stage a prepared kit as Kubernetes ConfigMaps and Secrets and mount them through the generated Helm chart; the workspace PVC stays mounted for writable runtime state.
Run
nvflare deploy k8s stageafternvflare deploy prepare; add--kubectl ocfor OpenShift and runnvflare deploy k8s unstageafter Helm uninstall.See Running FLARE in Kubernetes and the OpenShift / multicloud examples.
Hugging Face Client API — federate an existing
Traineror TRLSFTTrainerthroughflare.patch(trainer).FLARE owns round exchange, global-weight loading, local-budget enforcement, rank-0 communication, checkpoint continuity, and metric reporting.
See HuggingFace Client API and the Hello Hugging Face example.
Client API Attach and Recipe updates — Attach mode lets an independently started, externally owned trainer connect to FLARE without transferring process ownership to NVFlare (unlike
in_process/external_process, where FLARE launches and owns the trainer process).See Client API Attach Mode and the example.
Recipe updates add the concrete PyTorch FedBPT entry point, expose
key_metric_modefor FedAvg recipes, and improve PyTorch workflow support for FedProx, SCAFFOLD, Swarm, and model-selection behavior.
Compatibility and Migration Notes
Deployment and Security — changes every Docker, Slurm, Kubernetes, or POC deployment is likely to hit:
Job-process bootstrap credentials move off the command line. Launchers deliver them through the job process environment instead (a per-job Kubernetes Secret via
env[].valueFrom.secretKeyRef).No fallback: Docker/Kubernetes job images must run NVFlare 2.9 or newer, or they fail immediately at argument parsing when launched by a 2.9 CP/SP. The CLI path is retained, so an older parent launching a newer job image is unaffected.
A custom launcher that renders worker commands from
generate_client_command/generate_server_commandand implementslaunch_jobdirectly must also exportget_credential_env(job_args)into the child environment.Launcher Kubernetes RBAC now needs the
patchanddeleteverbs on Secrets (already in the generated Helm role templates).
Internal CellNet TCP links default to mTLS. Docker, Slurm, Kubernetes, and Network Attach deployments now default to mutual TLS between a parent and its job processes; a site that intentionally runs without it needs an explicit clear-transport opt-out. See Security Hardening above for the full list of hardening changes.
Kubernetes and Slurm sites use the participant certificate in both TLS roles, which requires a certificate allowing both
clientAuthandserverAuth. A startup kit with a role-restricted certificate — for example, from NVFlare 2.8 distributed provisioning, which issues onlyclientAuthor onlyserverAuth— must be re-provisioned before using the Kubernetes or Slurm launcher; unrestricted (no-EKU) certificates remain compatible. After re-provisioning, Kubernetes sites rerunnvflare deploy prepareand Slurm sites rebuild the runtime workspace.
Docker job-controlled launcher options are now restricted. Jobs may control only
image,python_path,entrypoint,num_of_gpus, andshm_sizethrough their launcher metadata; selectingimage,python_path, orentrypointrequires BYOC authorization at each receiving site.Previously job-controlled Docker SDK options such as
ipc_modeanddevice_requestsare now site-owned, configured throughdefault_job_container_kwargsor a study’sdocker_kwargs; launcher-owned options such as mounts and networks remain fixed.New jobs with unsupported options are rejected at submission; jobs stored before an upgrade are checked again and can fail at launch until their metadata is migrated.
Portable job resource fields are now reserved. The flat
resource_specnamesnum_of_gpus,num_of_cpus, andmemorymust use the documented portable types. Custom resource managers that previously interpreted these names differently must migrate to the portable types or rename their custom fields. Legacy nested resource specifications without@defaultremain unchanged.``poc start``/``poc stop`` preserve every repeated flag. Earlier versions silently kept only the last
-p/--serviceor-ex/--excludevalue.poc stopnow also honors participant exclusions, and now waits for targeted and exclusion-based shutdowns to complete before returningstatus: stopped(use--no-waitfor fire-and-forget). A barepoc startcontinues to start the server and clients without an admin console.
Client API and Recipes — affects any job built on the Client API or recipe framework:
Unified Client API execution paths.
ClientAPIExecutorconsolidates NVFlare’s trainer-process ownership patterns behind one Client API executor; jobs generated with FLARE 2.9 require a client runtime that provides it and are not runnable on older client runtimes.in_processreplaces the previousInProcessClientAPIExecutor.external_processreplaces the formerClientAPILauncherExecutorstack (LauncherExecutor,SubprocessLauncher,TaskExchanger,FlareAgent,BaseScriptRunner,ExternalConfigurator, and thePipe/PipeHandlerimplementations includingFilePipeandCellPipe) for trainers launched and owned by NVFlare: the launched trainer creates its own Cell, with a prescribed FQCN from a typed, owner-only bootstrap file, and connects to the client job’s Cell over an authenticated, liveness-checked session.attachprovides a standard Client API migration path for the independently managed trainer pattern served byIPCExchangerandIPCAgent. It preserves the server-facing trust boundary and CP-routed topology of that path while adding an explicit session protocol, Attach profiles, andflare.receive()/flare.send()integration.Custom parameter transformations (formerly
ParamsConverterand the framework-specific converter components) belong in trainer code aroundflare.receive()/flare.send(); common functions remain available innvflare.client.converter_utils.Recipe-level
pipe_typeandpipe_root_pathoptions are also removed; transport is selected through site communication configuration. The F3FileDriverremains available as schemeshared-filefor an attached trainer; a launched external-process trainer instead requires a clear TCP listener bound to loopback.ScriptRunnerselectsClientAPIExecutor(in_process)by default, orClientAPIExecutor(external_process)whenlaunch_external_process=True, and no longer performs a build-time PyTorch or TensorFlow import check. A client app may contain only oneClientAPIExecutor; configurations that previously added multiple script runners to one site must combine the scripts behind one entry point and dispatch on the Client API task name.
FedProx recipe rename. Recipe discovery now exposes the concrete PyTorch
FedProxRecipeasfedprox-ptand no longer advertises thefedprox-tfmanual pattern as a concrete recipe. TensorFlow clients can still combine a FedAvg recipe withTFFedProxLossexplicitly.
Streaming and Collaboration API — narrower audience: sites moving large models, and early adopters of the Technical Preview Collaboration API:
Streaming transport defaults are larger. The sender’s default streaming window / ACK interval move from 2.8’s 16 MiB / 4 MiB to 64 MiB / 16 MiB — set via the
streaming_window_size/streaming_ack_intervalkeys incomm_config.yml(seedev_tools/f3/comm_config.ymlfor the shipped values);TCP_NODELAYis now on by default, reducing request/ACK latency.A 2.8 receiver has a fixed out-of-sequence tolerance of 16 chunks and cannot derive a larger one from a peer’s advertised window; a 2.9 sender’s 64 MiB default window can exceed that under load and abort the stream. While any 2.8 peers are still in the fleet, set
streaming_window_size: 16Mon 2.9 senders that talk to them.
Collab calls carry a versioned authorization envelope. Calls are accepted only from authenticated participants in the same job. The Collaboration API is new in 2.9.0 (there is no earlier Collab implementation to be compatible with); every site running a Collab job must run 2.9.0 or newer.
Framework and Workflow-Specific Changes — scoped to Lightning, CCWF, and Swarm Learning users:
Patched Lightning clients report real per-round steps.
NUM_STEPS_CURRENT_ROUNDis now the actual per-round change intrainer.global_stepinstead oftrainer.estimated_stepping_batches, correcting cumulative aggregation over-weighting in later rounds whenupdate_fit_loop=True.global_stepcounts steps across all optimizers, so a multi-optimizer FedAvg client reports their combined step count unless it suppliesNUM_STEPS_CURRENT_ROUNDexplicitly; explicit client metadata is still preserved.
Patched Lightning clients transmit pre-fit validation metrics. A metric captured by an explicit
trainer.validate()call beforetrainer.fit()is now transmitted regardless oftrain_with_evaluation; sanity-check and in-fit validation metrics are still not transmitted as global-model scores. This enablesIntimeModelSelectorand best-global-model persistence for recipe-based Lightning jobs.train_with_evaluation=Truestill requires validation metrics; otherwise metrics remain optional, butFalseno longer suppresses metrics from an explicit pre-fit validation.An application that must keep such metrics local should omit the explicit pre-fit validation; if it still requires one locally, a custom task-result filter must remove
MetaKey.INITIAL_METRICSbefore the result reaches the server.
``SimpleIntimeModelSelector`` (CCWF) now handles dict-valued metrics. It selects a scalar
key_metric(defaultval_accuracy) from dict-valuedINITIAL_METRICSpayloads instead of failing with a loggedTypeErrorthat silently disabled best-model selection.Swarm jobs whose metric dicts contain the configured key change from selection-inert to active best-global-model tracking without a config change; dict payloads lacking the key, and non-numeric values, are skipped with a warning, so configure
key_metricto match the reported metric name.A new
negate_key_metricargument supports lower-is-better metrics such as losses.
Swarm can combine tensor streaming with disk offload. Set
aggregation_format=ExchangeFormat.PYTORCHandenable_tensor_disk_offload=TrueonSwarmLearningRecipe; the same offload flag is available onSwarmClientConfigfor Job API users.``SwarmLearningRecipe`` now configures best-model selection by default. Use
key_metricto select a dictionary-valued validation metric andkey_metric_mode="min"for lower-is-better metrics.Clients must report a pre-training validation metric with the configured name for selection to occur; jobs without that metric continue to persist the last global model but do not create
best_FL_global_model.pt. Setkey_metric=Noneto opt out and preserve the pre-2.9 last-model-only behavior.Selection skips round 0, so a one-round job does not create a best-model checkpoint. With
key_metric_mode="min", Swarm best-metric logs and records expose the negated comparison value (for example, a loss of 2.31 is shown as -2.31).client_config_overridescan no longer replacemodel_selector: migrate the former{"model_selector": None}opt-out tokey_metric=None, and useBaseSwarmLearningRecipewith an explicitSwarmClientConfigfor a custom selector.
Auto-FL — Auto-FL optimization is a new agent-directed campaign feature in 2.9.0 (see Agent Skills above), not an update to an existing example; the following are deep, edge-case-heavy compatibility notes scoped to its campaign admission behavior:
Auto-FL campaigns now honor the job’s native metric direction (
key_metric_modeor a matching same-metricstop_cond) instead of assuming maximization, so raw lower-is-better objectives no longer need to be negated. Campaign admission also fails closed in the following cases:An obvious lower-is-better metric such as
val_lossthat relies only on NVFlare’s implicitmaxdefault is rejected until the job declareskey_metric_mode="min".A job passing a custom
model_selectoris rejected, because that component supersedeskey_metric_modeand its selection direction can’t be imported deterministically. Remove the custom selector and expose its criterion as a declaredkey_metricwithkey_metric_modebefore initializing a campaign.A requested metric that differs from the job’s key metric is rejected unless
mutation_schema.yamldeclares the requested and optimization metric bridge. A declared bridge cannot override an unresolved native job metric.Job constructor calls that pass positional arguments,
*args, or**kwargsalso fail closed, since dynamic arguments could hide the metric, direction, or fixed training budget; rewrite the call with keyword-only arguments and no splats.SimEnvcalls must pin a positive explicitnum_clients, or expose a non-empty staticclientslist whose length is pinned.Experimental legacy minimization campaigns without direction provenance must be re-initialized in a fresh workspace.
External-process Client API: an accepted lazy result can’t be withdrawn. Losing a trainer after its lazy result envelope has been accepted now fails the run as
EXECUTION_EXCEPTION, even if a controller’smin_responsesthreshold could otherwise tolerate a missing client — the accepted envelope may already have exposed references to downstream consumers. An explicit job abort that wins the terminal-state race remainsABORTED.
Deprecations
HUB support removed. The deprecated FL HUB feature and its remaining runtime, documentation, and test surface are removed, including the
nvflare.app_common.hubimplementation and its component registry entries.