added docker compose for running local arma network run with grafana - #1203
moran-abilea wants to merge 2 commits into
Conversation
1a83e2c to
6b0df3a
Compare
b105104 to
184028f
Compare
There was a problem hiding this comment.
🟡 Changes recommended
The documented build and startup flow fails, and the dashboard contains a non-working metric query.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
Adds a local Docker Compose example running Arma with Prometheus, Grafana, and a configurable load test.
Changes:
- Adds setup, cleanup, and Docker build scripts.
- Adds Prometheus/Grafana provisioning and monitoring dashboard.
- Documents local usage and configuration.
File summaries
| File | Summary |
|---|---|
node/examples/grafana/scripts/run_sample.sh |
Generates configuration and starts the sample network; composed node paths currently prevent services from starting. |
node/examples/grafana/scripts/clean_sample.sh |
Stops containers and removes generated data. |
node/examples/grafana/scripts/build_docker.sh |
Builds the Arma image; the documented build command and image tag are currently incorrect. |
node/examples/grafana/README.md |
Documents setup and operation; its Grafana volume guidance does not match the Compose configuration. |
node/examples/grafana/grafana/datasource.yaml |
Provisions Prometheus as Grafana’s data source. |
node/examples/grafana/grafana/dashboards.yaml |
Configures dashboard provisioning. |
node/examples/grafana/grafana/arma-dashboard.json |
Defines monitoring panels; one batcher latency query uses the wrong metric namespace. |
node/examples/grafana/compose.yaml |
Defines Prometheus, Grafana, and submitter services; Grafana data is not persisted. |
Review details
Suppressed comments (2)
node/examples/grafana/README.md:95
- The README says edits live in a Grafana volume, but this compose file declares no Grafana data volume; only the read-only provisioning and dashboard bind mounts are present. Dashboard edits therefore live in the container's ephemeral filesystem and are lost when the container is recreated or removed by cleanup. Please either add a persistent volume and explain its lifecycle, or document that edits must be exported before teardown without calling it a volume.
Edits live in Grafana's volume, which `clean_sample.sh` removes, so export the JSON over `grafana/arma-dashboard.json` to keep one.
node/examples/grafana/compose.yaml:30
- The Grafana data directory is not persisted: these mounts only provide read-only provisioning files, and no volume is mounted at
/var/lib/grafana. Consequently dashboard edits made through the UI are lost when the container is recreated, contrary to the README's claim that edits live in Grafana's volume. Add a named Grafana data volume (and keep the cleanup script removing it) if UI edits are meant to survive a restart.
- /tmp/arma-sample/grafana/provisioning:/etc/grafana/provisioning:ro
- /tmp/arma-sample/grafana/dashboards:/etc/grafana/dashboards:ro
- Files reviewed: 7/8 changed files
- Comments generated: 3
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| cd "$(dirname "${BASH_SOURCE[0]}")/../.." | ||
| bash ./scripts/build_docker.sh |
There was a problem hiding this comment.
fixed as suggested,
also not written here in the review note but also fixed as suggested: added correct links in the readme.md
| chmod a+x "${SAMPLE_DIR}" | ||
| chmod -R a+rX "${SAMPLE_DIR}/grafana" "${SAMPLE_DIR}/prometheus.yml" | ||
|
|
||
| "${COMPOSE[@]}" up -d |
There was a problem hiding this comment.
checked it, it's wrong
nothing to fix here
| }, | ||
| "spec": { | ||
| "editorMode": "code", | ||
| "expr": "sum by (party_id, shard_id) (\r\n batch_ledger_append_latency_seconds_sum\r\n)\r\n/\r\nsum by (party_id, shard_id) (\r\n batch_ledger_append_latency_seconds_count\r\n)", |
There was a problem hiding this comment.
checked it, it's wrong
nothing to fix here
eaf18a1 to
0acdc3a
Compare
|
|
||
| # The same "arma" image the other examples use. | ||
| cd "$(dirname "${BASH_SOURCE[0]}")/../.." | ||
| bash ./scripts/build_docker.sh |
There was a problem hiding this comment.
I don't think this wrapper file is needed, you can use the existing node/examples/scripts/build_docker.sh directly
There was a problem hiding this comment.
fixed as suggested
| # The generator leaves the operations port unset, which means a random one, and binds it | ||
| # to localhost. Prometheus scrapes it from another container and needs to know it in | ||
| # advance. Every node has its own container, so one port serves them all. |
There was a problem hiding this comment.
maybe this comment can be shorter, for example Expose a fixed operations port so Prometheus can scrape the nodes.
There was a problem hiding this comment.
fixed as suggested
| mkdir -p "${SAMPLE_DIR}/grafana/provisioning/datasources" \ | ||
| "${SAMPLE_DIR}/grafana/provisioning/dashboards" \ | ||
| "${SAMPLE_DIR}/grafana/dashboards" | ||
| cp "${EXAMPLE_DIR}/grafana/datasource.yaml" "${SAMPLE_DIR}/grafana/provisioning/datasources/" | ||
| cp "${EXAMPLE_DIR}/grafana/dashboards.yaml" "${SAMPLE_DIR}/grafana/provisioning/dashboards/" | ||
| cp "${EXAMPLE_DIR}/grafana/arma-dashboard.json" "${SAMPLE_DIR}/grafana/dashboards/" |
There was a problem hiding this comment.
Do we need to copy the grafana files into /tmp? They can be mounted directly from compose.yaml instead
There was a problem hiding this comment.
I tried mounting these files directly from the repository, and it simplifies the script. However, it makes the example dependent on the checkout permissions. With umask 027, the files are 640 and Grafana fails to read the bind mounts.
We could chmod the repository files before starting Compose, but I'd prefer not to modify permissions in the user's checkout. An init container also works, but adds a service and named volumes just to stage three files.
So I kept the small staging step in /tmp, and added a comment explaining that it intentionally normalizes permissions for the container.
| # Scrape every node of the network. The hosts come from the deployment, so to change the | ||
| # number of parties or shards edit ../config/example-deployment.yaml and ../compose.yaml. |
There was a problem hiding this comment.
I don't think this comment is needed
There was a problem hiding this comment.
removed entierly
| ``` | ||
|
|
||
| The number of parties and shards comes from the network itself, so changing it means | ||
| editing [../config/example-deployment.yaml] and |
There was a problem hiding this comment.
this looks like a link, but it doesn't have a target
There was a problem hiding this comment.
alread fixed, with the rest of the links
| ### Working on a New Metric Methodology | ||
|
|
||
| 1. Add the metric in `node/<role>/metrics.go`, keeping `party_id` and `shard_id` among its | ||
| `LabelNames` as the existing metrics do. | ||
| Its Prometheus name is `<namespace>_<name>`. | ||
| 2. Rebuild the image and start again: `build_docker.sh`, `clean_sample.sh`, `run_sample.sh`. | ||
| 3. Confirm it is exported, in the Prometheus UI or straight from a node: | ||
| `docker exec arma-grafana-prometheus-1 wget -qO- http://router.p1:8080/metrics | grep <name>` | ||
| 4. Add a panel in Grafana, selecting `${DS_PROMETHEUS}` as its data source rather than the | ||
| Prometheus data source itself, so the dashboard stays independent of this setup. | ||
| 5. Export it over `grafana/arma-dashboard.json` and document the metric in [docs/monitoring/metrics.md]. |
There was a problem hiding this comment.
I think this section is a bit out of scope here.
There was a problem hiding this comment.
section removed entierly
8d56e83 to
967fce4
Compare
Signed-off-by: Moran Abilea <moran.abilea@ibm.com>
Signed-off-by: Moran Abilea <moran.abilea@ibm.com>
967fce4 to
2af5d4f
Compare
related to #449
Running Arma with Prometheus and Grafana inside docker compose
it runs an Arma sample network together with Prometheus and
Grafana, and watch the metrics of a load test on a dashboard while it runs.
By default it runs a network of four parties with two shards — each party has a
router, two batchers, a consensus and an assembler node — and drives it with a five
minute load test. The number of parties, the number of shards, the duration and the
load are all configurable.
Everything runs as the user that invokes the scripts.