nvflare.job_config.script_runner module

class ScriptRunner(script: str, script_args: str | list[str] = '', launch_external_process: bool = False, command: str | list[str] = 'python3 -u', framework: FrameworkType = FrameworkType.PYTORCH, server_expected_format: ExchangeFormat = ExchangeFormat.NUMPY, params_transfer_type: TransferType = TransferType.FULL, launch_once: bool = True, launch_timeout: float | None = 300.0, shutdown_timeout: float = 0.0, memory_gc_rounds: int = 0, cuda_empty_cache: bool = False, execution_mode: str | None = None)[source]

Bases: object

Adds a Client API training script to a FedJob.

Transport is selected by the site’s Cell driver configuration. The runner only selects whether the trainer runs in the Client Job process or in a process owned and launched by NVFlare.

Initializes the runner.

Parameters:
  • script – Training script path.

  • script_args – Arguments appended to the script. Pre-tokenized argv preserves exact argument boundaries for external processes.

  • launch_external_process – Select external_process when execution_mode is omitted; otherwise select in_process.

  • command – Command prepended to the script in external_process mode.

  • framework – Trainer-native parameter representation.

  • server_expected_format – Parameter representation expected by the server.

  • params_transfer_type – Whether the trainer returns FULL parameters or a DIFF.

  • launch_once – Launch once per job or once per task in external_process mode.

  • launch_timeout – Maximum time for the external trainer to initialize and connect.

  • shutdown_timeout – External-process orderly-exit wait. This also feeds several accepted-result cleanup bounds in ClientAPIExecutor. The default zero skips direct orderly-exit and finalize-gate waits, but maps to 30 seconds for accepted- source disconnect and post-settlement group-exit waits and feeds the fixed settled- reaper budget; see ClientAPIExecutor.shutdown_timeout for all roles.

  • memory_gc_rounds – Force memory cleanup every N rounds; zero disables it.

  • cuda_empty_cache – Empty the CUDA cache during configured memory cleanup.

  • execution_mode – Optional explicit in_process or external_process mode. Use ClientAPIExecutor directly for an independently managed trainer in attach mode.

add_to_fed_job(job: FedJob, ctx, **kwargs)[source]

Adds the configured ClientAPIExecutor and script resource to the job.