Upgrade Guide r3.5.0
This guide is for operators who upgrade an OctoMesh installation from r3.4.x to r3.5.0. What is new is summarised in the Release Notes r3.5.0.
The Communication Controller migrates every tenant from System.Communication 3.x to 4.x when it starts. There is no downgrade path: r3.4.x code cannot read 4.x data. The only way back is to restore the database backup taken before the upgrade, and everything written after the migration is lost. Do not start without a verified backup.
What changes
| Area | Change | Action needed |
|---|---|---|
| CK models | System.Communication 3.x → 4.6.0, System.Ai 3.x → 4.3.0 | none, the migration runs automatically |
| Data | Pool entities become DeploymentSite; Manages/ManagedBy edges become Hosts/HostedBy | verify after the upgrade |
| Kubernetes | CRD CommunicationPool (v1alpha1) is replaced by DeploymentSite (v1); no conversion | install the new CRD, redeploy every site, delete the old resources and CRD |
| Operator Helm values | autoManagePools, poolNamespace, defaultPoolName renamed | update every values file and overlay |
| Core Helm values | new hub authorization and SECRET key ring values | none; they are inactive until configured |
| Edge operators | operator ↔ controller protocol renamed with the CRD | upgrade in the same maintenance window as the controller |
| REST / octo-cli / MCP | /v1/pool → /v1/deploymentsite, *Pool* verbs → *DeploymentSite* | update scripts and integrations; use the octo-cli version that matches the cluster |
| Blueprints | System.Communication-[x,4.0) ranges cannot be satisfied any more | update each affected blueprint to its new major version (see Blueprints and CK models) |
| Blueprint installation | dependencies resolve to the highest matching version across all catalogs; installing never downgrades a dependency | check scripts that relied on the old behaviour |
| Adapters | SDK built for 4.x; OCTO_ADAPTER__TENANTID deprecated; TLS validation enforced on the hub connection | re-release every externally built adapter, check certificates |
| Pipelines | cron executions report the trigger type Scheduled | update dashboards that count cron runs as Event |
| Metrics / alerting | octo.workload.kind="pool" → "deployment_site" / "adapter_pool" | migrate dashboards and alert rules in the same window |
| Refinery Studio | Pools become Deployment Sites, no redirect for communication/pools | roll the Studio out last |
| Adapter pools and leasing | new, off by default | nothing; do not enable it during the upgrade |
Before you upgrade
-
Plan a maintenance window that covers the core services, the operators (central and edge), the adapters and the blueprint updates of the installation. Adapters keep running during the window, but deployments and workload changes are not possible between the controller rollout and the operator rollout.
-
Stop automatic deployments. Pause every pipeline that deploys core services, operators, adapter charts or apps to the target cluster automatically — including pipelines that are triggered by builds of the main branch and pipelines that release adapters or apps on their own release line. A partial rollout, in which some services already run r3.5.0 and others still run r3.4.x, leaves tenants with unresolvable CK models. Keep them paused until the soak period after the upgrade has ended.
-
Take a backup inside the cluster of the system database (
octosystem), of every tenant database and of the job databases (Hangfire). Run it as a job in the cluster (for example withmongodump --archive --gzipin a backup pod) and copy it out of the cluster afterwards. A dump streamed to a workstation throughkubectl execor a port-forward does not reliably carry collections of several gigabytes. Check that the backup contains the large collections, in particular the GridFS collectionsfs.filesandfs.chunksofoctosystemand the event collections of the tenants. Back up the custom resources and the Helm values as well:kubectl get communicationpools.octo-mesh.meshmakers.io -A -o yaml > communicationpools-backup.yamlhelm get values <release> -n <namespace> -o yaml > <release>-values.yamlThe stream data store (CrateDB) is not touched by the migration; a backup is optional.
-
Record the current state of every tenant so that you can compare it after the migration:
octo-cli -c LibraryStatusocto-cli -c ListBlueprintInstallationsNote the installed
System.Communicationversion and count the entities that will be migrated (see Verify the migration). Record theHelmRepositoryassociations of every adapter as well (see Helm repository associations). -
Check the tenants'
System.Communicationversion. r3.5.0 migrates tenants on 3.35.0 up to 3.41.0. A tenant on 3.42.0 is refused (see Data migration). -
Check the blueprints. Every installed blueprint that depends on
System.Communicationwith an upper bound below 4.0 (for exampleSystem.Communication-[3.22,4.0)) needs its new major version. Make sure those versions are published to the catalog before you upgrade (see Blueprints and CK models). -
Find all old Helm keys. Helm silently ignores unknown values. An operator that still receives
autoManagePoolsfalls back to edge mode and deploys nothing. Search every values file and overlay:grep -rnE 'autoManagePools|poolNamespace|defaultPoolName|communicationpools|CommunicationPool' <your-deployment-repo> -
Check certificates of edge sites. From r3.5.0 adapters validate the TLS certificates of the platform services on every connection, including the hub connection (see TLS certificate validation). An adapter at an edge site that connects through a private or self-signed certificate, or through a proxy with TLS inspection, needs the issuing CA in its trust store.
-
Check the edge devices. CRDs are cluster-wide. If one edge device runs operators for more than one installation, its CRD and operator upgrade affects all of them; plan that device into the window of the installation that is upgraded last, or move the other installation's site first.
-
Prepare octo-cli in both versions. Use the r3.4.x octo-cli against clusters that still run r3.4.x and the r3.5.0 octo-cli against upgraded clusters (see API, octo-cli and MCP).
Upgrade order
Run the steps in this order and wait for each workload to become ready before you continue.
-
CRDs: install the CRD chart of r3.5.0. It adds
deploymentsites.octo-mesh.meshmakers.io. The oldcommunicationpoolsCRD carrieshelm.sh/resource-policy: keepand stays in place; leave it for now. -
Old operator: scale the r3.4.x Communication Operator to zero replicas so that it does not act on the old custom resources while the controller migrates.
-
Core services in one go: roll out all core services to r3.5.0 — Identity first (it imports the
Systemmodel), then Asset Repository, Platform, Bot and finally the Communication Controller. The Communication Controller importsSystem.Communication4.x and migrates every tenant. No core service may stay on an r3.4.x image.Bot before the controllerThe Bot service must be ready (pod ready and its Hangfire server registered) before the Communication Controller starts. Otherwise the controller migrates the data but fails to start most tenants with
Failed to start deferred tenant, and they do not recover by themselves. A Helm upgrade of the core chart rolls all services at the same time, so check the controller log after the rollout. If the message appears, or if the Bot restarted while the controller was starting, restart the Communication Controller once the Bot is ready. -
AI, Reporting and further services: roll out the charts of the remaining platform services (AI, Reporting, Office, MCP) to r3.5.0.
-
Operators immediately after the controller: roll out the r3.5.0 Communication Operator with the renamed Helm values (see Operator and Helm values), centrally and on every edge cluster. An r3.4.x operator cannot talk to an r3.5.0 controller.
-
Redeploy the Deployment Sites per tenant. There is no conversion of the custom resources, and the operator does not recreate them by itself after the upgrade: the controller keeps the list of deployed sites in memory, so after its restart a freshly started operator registers with zero sites, even with
autoManageDeploymentSites: true. Redeploy every site of every tenant:octo-cli -c GetDeploymentSitesocto-cli -c DeployDeploymentSite -id <deploymentSiteRtId>This creates the
DeploymentSiteresources; running adapter pods are not restarted. -
Adapters: roll out the r3.5.0 mesh adapter chart and the r3.5-line releases of all externally built adapters (see Adapters).
-
Per tenant: refresh the catalogs, check the CK models and clear the caches. Run the tenants one after the other, not in parallel:
octo-cli -c RefreshCatalogsocto-cli -c RefreshBlueprintCatalogsocto-cli -c LibraryStatusocto-cli -c FixAll -w -yocto-cli -c ClearCacheThen verify the migration and update the blueprints to their new majors.
-
Refinery Studio last, after all back-end services run r3.5.0.
-
Remove the old custom resources and the old CRD once every site is deployed again. The old resources carry a finalizer of the r3.4.x operator, which no longer runs; remove it before you delete them. Deleting the CRD deletes every remaining
CommunicationPoolresource in all namespaces, so make sure no other installation in the same cluster still uses it:kubectl get communicationpools.octo-mesh.meshmakers.io -Akubectl patch communicationpools.octo-mesh.meshmakers.io <name> -n <namespace> \--type merge -p '{"metadata":{"finalizers":[]}}'kubectl delete communicationpools.octo-mesh.meshmakers.io <name> -n <namespace>kubectl delete crd communicationpools.octo-mesh.meshmakers.io -
Observability: switch dashboards and alert rules from
octo.workload.kind="pool"todeployment_site(andadapter_pool) in the same window. New series can show<no value>for the first few evaluations. -
Resume automatic deployments only after a soak period (we recommend 24 hours) without new error classes.
Data migration
The migration is part of the System.Communication 4.x model and runs automatically when the Communication Controller imports it into a tenant.
-
Entry points: every published 3.x version from 3.35.0 to 3.41.0 (3.35.0, 3.36.0, 3.37.x, 3.38.0, 3.39.x, 3.40.0, 3.41.0) migrates to 4.6.0 with the same script; the 4.x versions after 4.0.0 are additive.
-
Not an entry point: 3.42.0.
System.Communication3.42.0 stores some attributes as encrypted SECRET values. r3.5.0 deliberately has no migration entry for it, because it would rename encrypted values onto plain-text attributes. A tenant on 3.42.0 is refused (see the safety net below) and stays on 3.x; it is migrated by a later r3.5.x release whose model contains the 3.42.0 entry point. -
What is converted: the CK type of every
Poolentity changes toDeploymentSite(rtId, well-known name and attribute values are kept); everyManagesedge becomes aHostsedge; adapters keep their deployment site. -
What is not converted: the Kubernetes custom resources (see above) and runtime models or seed files that you maintain yourself. Files that still use
System.Communication/Poolor the roleManagesfail to import on 4.x and need updating. -
Safety net: an upgrade that would cross a major version without a migration entry is refused. The CK import then reports:
A schema-only bridge across a major version is refused because it would skip the data migration.The tenant stays on its 3.x data. Do not work around it; contact support.
Verify the migration
For every tenant, compare before and after:
| Check | Expected after the upgrade |
|---|---|
LibraryStatus | System.Communication 4.6.0, all models resolved, no ResolveFailed, nothing that needs action |
Entities of type System.Communication/Pool | 0 |
Entities of type System.Communication/DeploymentSite | the number of Pool entities before |
Edges with role Manages | 0 |
Edges with role Hosts | the number of Manages edges before |
| Adapters | the same number as before, each with its deployment site |
Deployment Sites in GetDeploymentSites | all sites, Online after the operator reconnected |
DeploymentSite resources in Kubernetes | one per site, status Registered |
| Controller log | no Failed to start deferred tenant, no System tenant database does not exist |
Helm repository associations
The r3.5.0 Communication Controller re-applies the seed of the System.Communication blueprint when it starts. An adapter that was associated with two HelmRepository entities keeps only one of them; the second association (typically the one to a release chart repository) is removed. Compare the associations you recorded before the upgrade and restore the missing ones where the adapter needs them, before you roll out new adapter charts.
Blueprints and CK models
r3.5.0 is published together with new major versions of every blueprint and CK model that depends on System.Communication. The 1.x (and other pre-4.x) lines stay in the catalogs in parallel, so installations that still run r3.4.x keep using them; do not install them on an r3.5.0 installation.
| Model or blueprint | Version for r3.5.0 |
|---|---|
System.Communication (CK, imported by the controller) | 4.6.0 |
System.Ai (CK, imported by the AI service) | 4.3.0 — credentials are SECRET attributes |
Loxone (CK) | 5.1.0 — the Miniserver password is a SECRET attribute. 5.0.0 is the same model with a plain-text password; tenants on 5.0.0 can update to 5.1.0 |
MeshmakersAccounting / .Tesla / .Host | 2.2.0 / 2.1.0 / 2.0.0 |
EnergyCommunity.Base / .Billing / .Simulation | 2.11.1 / 2.8.1 / 2.8.2 |
EnergyCommunity.App / .EdaIntegration | 2.1.0 / 2.11.0 |
OneTimeTicket.Release / .MainLatest | 2.0.0 |
FdaSeen.Base | 2.0.0 |
ZenonDynprop.MainLatest | 2.0.0 |
Samples.PipelineBasics / .Photovoltaics / .Simulator.EnergyCommunity | 2.0.0 / 2.0.1 / 2.0.0 |
Office.ExcelImport | 2.0.0 |
SmartMeterInsights.Base | 2.0.0 |
FamilyOs.Release / .MainLatest | 2.0.0 |
Update each installed blueprint to the highest version of its new major, with merge mode. For energy communities update in the order Base → Billing / EdaIntegration → Simulation → App. Preview the update first:
octo-cli -c RefreshBlueprintCatalogs
octo-cli -c PreviewBlueprintUpdate -tv MeshmakersAccounting-2.2.0
octo-cli -c UpdateBlueprint -tv MeshmakersAccounting-2.2.0 -m Merge
Install CK models that are not part of a blueprint, such as Loxone, from the catalog and check the result with LibraryStatus afterwards (-w can report success before the model is resolved):
octo-cli -c ImportFromCatalog -cn <catalogName> -m Loxone-5.1.0 -w
octo-cli -c LibraryStatus
Changed dependency resolution
- Highest matching version across all catalogs. A blueprint dependency range now resolves to the highest version that satisfies it in any readable catalog. Before, the first catalog with any matching version won, so a public catalog could hide a newer version from a private catalog.
- No downgrade. A dependency range is a minimum, not a target. When a tenant already runs a newer version of a dependency that satisfies every declared range, installing a blueprint keeps it — no seed import, no change to the installation record. When the installed version is outside a declared range, the installation fails instead of downgrading. Before, installing a blueprint could re-apply an older dependency's seed data and record the older version as installed.
- Blueprint catalogs can be switched off. The GitHub blueprint catalogs have an
IsEnabledoption like the CK catalogs. A disabled catalog is skipped by listing, search, dependency resolution, installation and refresh. The default istrue, which keeps the previous behaviour.
Rollback
There is no partial rollback. Rolling back only the images does not work, because r3.4.x code cannot read 4.x models. In a rehearsal with 24 databases the restore took about three to five minutes.
- Pause all automatic deployments to the cluster.
- Stop the r3.5.0 services.
- Drop
octosystem, every tenant database and the job databases, then restore them from the backup taken before the upgrade. Do not rely onmongorestore --dropalone: it drops only the collections that are in the archive, so collections created by r3.5.0 (for exampleRtEntity_SystemCommunicationLentAdapterPool) stay behind. - Roll the images, charts and operators back to the last r3.4.x release.
- Re-install the old CRD and restore the custom resources from
communicationpools-backup.yaml. - Clear the CK cache of every tenant (
octo-cli -c ClearCache) or restart the services.
Data written after the migration is lost. Decide on a rollback within the maintenance window; afterwards, fix forward.
Operator and Helm values
| r3.4.x value | r3.5.0 value |
|---|---|
operator.autoManagePools | operator.autoManageDeploymentSites |
operator.poolNamespace | operator.deploymentSiteNamespace |
operator.defaultPoolName | operator.defaultDeploymentSiteName |
The custom resource changes from
apiVersion: octo-mesh.meshmakers.io/v1alpha1
kind: CommunicationPool
spec:
tenantId: "meshtest"
poolRtId: "65d5c447b420da3fb12381bb"
to
apiVersion: octo-mesh.meshmakers.io/v1
kind: DeploymentSite
spec:
tenantId: "meshtest"
deploymentSiteRtId: "65d5c447b420da3fb12381bb"
See Deployment Sites — Installation for the complete resource.
Edge operators must be upgraded in the same window as the Communication Controller. An edge cluster that cannot be reached keeps its running workloads, but it cannot receive deployments until its operator runs r3.5.0.
New values that are inactive until configured
r3.5.0 adds Helm values that change nothing with their defaults. Configure them in a separate change after the upgrade.
| Value | Default | Effect when configured |
|---|---|---|
services.communication.hubAuthorization.adapterMode / operatorMode (core chart) | LogOnly | LogOnly only logs connections to the adapter and operator hubs that an enforcing controller would refuse and counts them in the metric octo.communication.hub.authorization.decisions (outcome=would_refuse). Enforce refuses them. Switch to Enforce per cluster only after the count has stayed at zero for a review period; operators need operator.authentication.* credentials first. Any other value fails the render |
operator.authentication.* (operator chart) | empty | client credentials with which the operator authenticates at the operator hub; required before operatorMode: Enforce |
secrets.communicationInstanceSecretKey, secrets.secretEncryptionKeys, secrets.secretEncryptionActiveKeyId (core chart) | empty | key ring for SECRET attribute values; empty means no key ring is rendered |
operator.clusterSecrets.instanceSecretKey and the matching key ring values (operator chart) | empty | key ring injected into workloads that receive cluster secrets; must match the core chart |
secrets.secretEncryptionRequired / operator.clusterSecrets.secretEncryptionRequired | false | true makes the render fail when the key ring is empty or not a valid 32-byte key; enable on central clusters once the key is verified, keep it off on edge operators without a key ring |
operator.adapterIgnoreCertificateValidation (operator chart) | false | turns off TLS certificate validation in every workload the operator deploys. For development clusters only; prefer secrets.rootCa (see TLS certificate validation) |
Adapters
Externally built adapters
Adapters compile the Construction Kit runtime and the communication SDK into their image. Every adapter that is not built from the platform release itself must be released again on the r3.5 line: finAPI, EDA, Loxone, SAP, weClapp (with the DILOS plug), Modbus (plug and socket), zenon, MQTT, the demo adapters and every customer-specific adapter. An adapter image built for r3.4.x keeps running after the core upgrade, but do not leave it on the old line: the next model change that it compiles in stops it, and an adapter that is restarted against a changed System model reports System tenant database does not exist.
- Roll out the re-released adapters after the core services of the cluster run r3.5.0. An r3.5-line adapter must not be deployed to a cluster that still runs r3.4.x.
- Release the finAPI adapter again in the same maintenance window as the core upgrade of each cluster, and build it against the released r3.5.0 SDK, not against packages of a development build.
- While some clusters still run r3.4.x, a hotfix for those clusters must be built against the r3.4.x SDK explicitly.
- Edge adapters that run as Windows services at a customer site (zenon, SAP) are installed manually; check their certificates (see below) before you hand them over.
Configuration key OCTO_ADAPTER__TENANTID
The adapter option that holds the adapter's tenant was renamed: AdapterOptions.TenantId → AdapterOptions.DedicatedTenantId, environment variable OCTO_ADAPTER__TENANTID → OCTO_ADAPTER__DEDICATEDTENANTID.
- In r3.5.0 the old key is still honoured and logs a warning at startup:
OCTO_ADAPTER__TENANTID is deprecated and will be removed in the next release. - When both keys are set, the new one wins.
- The old key is removed in the next release. Until then, adapter charts should render both keys with the same value; render only the new key once no image of the old line is deployed any more.
- Custom adapter code that reads
AdapterOptions.TenantIddoes not compile against the r3.5 SDK and must useDedicatedTenantId.
TLS certificate validation
The SDK now validates the server certificates of the platform services on all connections, including the SignalR hub connection between adapter and Communication Controller, which accepted every certificate in r3.4.x.
- Adapters that reach the platform through a public certificate are not affected.
- Adapters that reach it through a private CA, a self-signed certificate or a proxy with TLS inspection need that CA in their trust store. The adapter charts install it from the chart value
secrets.rootCawith an init container; adapters that run as Windows services need the CA in the Windows certificate store. IgnoreCertificateValidation=truedisables validation for test environments. It is refused whenASPNETCORE_ENVIRONMENTorDOTNET_ENVIRONMENTisProduction, and it logs a warning everywhere else. The operator valueoperator.adapterIgnoreCertificateValidationsets it for every workload of a cluster; leave itfalseoutside development clusters.
Datastore host check
An adapter with a configured datastore host name that does not resolve now stops at startup with an error that names the setting, instead of failing on its first execution.
API, octo-cli and MCP
| r3.4.x | r3.5.0 |
|---|---|
GET {tenantId}/v1/pool | GET {tenantId}/v1/deploymentsite |
POST {tenantId}/v1/pool/deploy?poolRtId=… | POST {tenantId}/v1/deploymentsite/deploy?deploymentSiteRtId=… |
POST {tenantId}/v1/pool/undeploy?poolRtId=… | POST {tenantId}/v1/deploymentsite/undeploy?deploymentSiteRtId=… |
POST {tenantId}/v1/pool/workloads/deploy / undeploy | POST {tenantId}/v1/deploymentsite/workloads/deploy / undeploy |
octo-cli -c GetPools | octo-cli -c GetDeploymentSites |
octo-cli -c DeployPool -id <poolRtId> | octo-cli -c DeployDeploymentSite -id <deploymentSiteRtId> |
octo-cli -c UndeployPool -id <poolRtId> | octo-cli -c UndeployDeploymentSite -id <deploymentSiteRtId> |
MCP get_pools | MCP get_deployment_sites |
MCP undeploy_pool | MCP undeploy_deployment_site |
Use the r3.5.0 version of octo-cli and of the MCP server against an r3.5.0 installation, and the r3.4.x version against an installation that has not been upgraded yet. This matters in particular for deployment automation that downloads the latest octo-cli: against an r3.4.x cluster, the r3.5.0 UpdateWorkloadChartVersion still succeeds, but the following DeployWorkload fails with HTTP 404 — the workload is left with the new chart version but not deployed. Pin the octo-cli version per cluster until every cluster runs r3.5.0.
GraphQL queries against SystemCommunicationPool must use the Deployment Site type instead.
Fixed: runtime query updates of to-one associations
Updating rows of a runtime query through GraphQL (runtimeQuery.update) could write association changes in the reversed direction, and replace to-many navigations, for association roles whose opposite side has the multiplicity ZeroOrOne. r3.5.0 decides the direction by the navigated side only. If you worked around the old behaviour, remove the workaround.
Observability
- The workload state metrics label sites with
octo.workload.kind="deployment_site"and adapter pools withocto.workload.kind="adapter_pool". There is no alias for"pool": rules and dashboards that filter on it stop matching. - Pipeline executions started by a cron schedule report the trigger type
Scheduled, on dedicated adapters as well as on adapter pools;Eventis reserved for real bus events. Dashboards and execution history filters that count cron runs underEventneed updating. - The Communication Controller counts hub authorization decisions in
octo.communication.hub.authorization.decisions. InLogOnlymode the series withoutcome=would_refuseshows which connections an enforcing controller would refuse. - Leasing adds metrics under
octo.lease.*andocto.pool.*. Alert rules for leasing are only useful on installations that enable it.
Adapter pools and leasing
Leasing is off on every tenant after the upgrade (leasingEnabled: false), and there are no adapter pools until you create one. Do not enable leasing as part of the upgrade. Enable it in a separate change, for one lending and one borrowing tenant first:
octo-cli -c GetCommunicationLifecycle
octo-cli -c SetCommunicationLifecycle -le true
A lease needs leasing to be enabled on both the lending and the borrowing tenant. Switching it off again holds the queue.
Known limitations
- Run the Communication Controller with one replica. Lease state lives in the controller process; persistent leases for several replicas are planned (AB#5878).
- Leasing is off by default on every tenant.
- Refinery Studio has no pages for adapter pools and leasing yet; they follow in a later Studio release. Until then use
octo-cli, the MCP tools or the REST API. - Adapter pools reject
ReceivesClusterSecretsand ingress configuration;ExecuteCSharp@1on a pool needs a memory limit of at least 1 Gi. - Leased executions cannot reveal SECRET attribute values and have no access to stream data (CrateDB).
- Changes to a pool reach borrowing tenants only after
POST {tenantId}/v1/adapterpool/mirrors/publish. - Tenants on
System.Communication3.42.0 are not migrated by r3.5.0 (see Data migration).
Related: Identity tenant API roles
r3.4.149 already changed the Identity tenant REST API: it requires the UserManagement or TenantManagement role (enforced by default) and adds GET users/directory for user pickers. If you upgrade from a release before r3.4.149, check that the users and clients that manage tenants hold one of these roles.