TownSuite FileSdConfigs
Prometheus service discovery

File-based service discovery for Prometheus, on autopilot.

A single self-contained executable that polls your services and continuously writes Prometheus file_sd_configs target files — for blackbox health checks, metrics scraping, OpenTelemetry, and DNS probing.

.NET self-contained
Cross-platform
Runs as a service
targets_v2.json
[
  {
    "targets": [
      "https://app.townsuite.com/healthz/live",
      "https://app.townsuite.com/healthz/ready"
    ],
    "labels": {
      "job": "app",
      "env": "prod"
    }
  }
]

How it works

1
Discover
Reads a service list, then each service's instances and endpoints — via SettingsV2.
2
Poll
Loops every DelayInSeconds, refreshing targets quietly in the background.
3
Write
Emits clean file_sd JSON for health checks, metrics, OpenTelemetry and DNS.
4
Scrape
Prometheus picks up the files via file_sd_configs — no restarts needed.

Output files

targets_v2.json
Health-check endpoints per instance — feed this to the blackbox exporter.
targets_prometheus_metrics.json
BaseUrl + PrometheusMetricsUrl for each discovered instance.
targets_dns.json
Scheme and host only, grouped per service, labelled job=dns_prober.
targets.json
Legacy v1 targets — every host crossed with AppendPaths, plus your static labels.
targets_open_telemetry_metrics.json
BaseUrl + OpenTelemetryUrl — only instances that declare the attribute.
Install

Build a single executable

Publish a self-contained, ready-to-run binary for your target platform. The result is one file — TownSuite.Prometheus.FileSdConfigs — that you drop onto the host.

macOS · zsh
dotnet publish -c Release -r osx-x64 \
  -p:PublishReadyToRun=true \
  --self-contained true \
  -p:PublishSingleFile=true \
  -p:EnableCompressionInSingleFile=true
Linux · bash
dotnet publish -c Release -r linux-x64 \
  -p:PublishReadyToRun=true \
  --self-contained true \
  -p:PublishSingleFile=true \
  -p:EnableCompressionInSingleFile=true
Windows · PowerShell
dotnet publish -c Release -r win-x64 `
  -p:PublishReadyToRun=true `
  --self-contained true `
  -p:PublishSingleFile=true `
  -p:EnableCompressionInSingleFile=true
Need a different architecture? Swap the runtime identifier (-r) — see the .NET RID catalog for the full list.
Configuration

appsettings.json

Drop an appsettings.json beside the executable. Define global behaviour under AppSettings, then list your sources under Settings or SettingsV2. The JSON your lookup endpoints must return is covered in Source JSON.

appsettings.json
{
  "AppSettings": {
    "DelayInSeconds": 1800,
    "UserAgent": "TownSuiteSD/1.0 ...",
    "OutputPath": "targets.json",
    "OutputPathV2": "targets_v2.json",
    "OutputPathPrometheusMetrics": "targets_prometheus_metrics.json",
    "OutputPathOpenTelemetry": "targets_open_telemetry_metrics.json",
    "OutputPathDns": "targets_dns.json",
    "HttpTimeoutInSeconds": 10,
    "SkipCertificateValidation": false,
    "MaxStaleCycles": 0
  },
  "Settings": [
    {
      "LookupUrl": "http host/path or filepath",
      "AuthHeader": "basic auth or bearer token",
      "AppendPaths": [ "/metrics", "/api/status" ],
      "Labels": { "job": "test", "env": "dev" },
      "IgnoreList": [ "https://example.townsuite.com" ]
    }
  ],
  "SettingsV2": [
    {
      "ServiceListUrl": "url/path that lists services",
      "ServiceDiscoverUrl": "url/path for service details",
      "AuthHeader": "basic auth or bearer token",
      "IgnoreList": [ "https://example.townsuite.com" ],
      "LowercaseLabels": true
    }
  ]
}

AppSettings reference

DelayInSeconds 1800
Seconds between refresh cycles. The poll loop sleeps this long between passes.
UserAgent string
User-Agent header sent on every lookup request. Omitted when left blank.
OutputPath targets.json
Path for the v1 file_sd targets output. Blank disables the v1 pass.
OutputPathV2 targets_v2.json
Path for the v2 (health-check) targets output. Blank disables it.
OutputPathPrometheusMetrics targets_..._metrics.json
Path for the Prometheus metrics endpoints output. Blank disables it.
OutputPathOpenTelemetry targets_..._otel.json
Path for the OpenTelemetry metrics endpoints output. Blank disables it.
OutputPathDns targets_dns.json
Path for the host-only DNS probing output. Blank disables it.
HttpTimeoutInSeconds 10
Per-request HTTP timeout applied to every lookup. Omitted or zero falls back to 100 seconds.
SkipCertificateValidation false
Accepts any TLS certificate when true. For internal hosts with self-signed certs only.
MaxStaleCycles 0
How many consecutive incomplete passes may keep serving the previous target files. Zero holds the last good files for as long as the lookups keep failing.

When the configuration is wrong

Each output is independent — a path left blank disables just that one and the rest still run. This is measured behaviour, and it applies to any of the output-path settings.

Key omitted, null or blank
That output is skipped for good — nothing written, no lookup run, other outputs unaffected.
A number, e.g. 12345
Configuration binds every value as a string, so this is a valid relative filename. You get a file literally named 12345.
An object or array
Binding fails and the process exits before the first pass.
Missing or unwritable directory
The pass throws, logs Crashing at critical and exits -1. Outputs written earlier in the pass keep their files; systemd or NSSM then restarts in a loop until the path is fixed.
SettingsV2 section missing
Startup fails — the section is required even when only v1 sources are used. Use an empty array for none.
SettingsV2 empty array
Every v2 output is written as a valid empty document.
Null entry, or no ServiceListUrl
The entry is logged and skipped and the pass is marked incomplete; the remaining entries still produce targets.
HttpTimeoutInSeconds omitted or 0
Falls back to 100 seconds instead of failing on a zero timeout.
Settings v1 sources

Point at a list of hosts and append metric paths. Best for static, known endpoints.

LookupUrl url or filepath listing base hosts
AuthHeader sent as the Authorization header
AppendPaths paths appended to every host
Labels static labels for each target
IgnoreList URLs to skip (exact match)
SettingsV2 dynamic discovery

Resolves a service list, then discovers details per service. Best for fleets that change.

ServiceListUrl url or filepath listing service keys
ServiceDiscoverUrl per-service detail lookup, key appended
AuthHeader sent as the Authorization header
IgnoreList URLs to skip (exact or prefix match)
LowercaseLabels trim and lowercase every label
Source JSON

What your endpoints return

This is the JSON your own inventory service returns — not appsettings.json. Property names are matched case-insensitively, so BaseUrl, baseUrl and baseurl are all accepted.

1 · ServiceListUrl → the service keys

GET https://inventory.example.com/listservices
[
  "Service1.Example",
  "Service2.Example",
  "Payments.Example"
]
Each entry is split on . and only the first segment is used, so Service1.Example becomes the key Service1. That key is appended directly to ServiceDiscoverUrl — which is why that setting normally ends in = or /. A ServiceListUrl that does not start with http is read from disk instead, in the same array format.

2 · ServiceDiscoverUrl{key} → the instances

GET https://inventory.example.com/discover?service=Service1
{
  "version": "1",
  "services": [
    {
      "name": "Service 1 — primary",
      "id": "3f6b2c1e-9a44-4f2f-9c3f-6e5b1c8d0a11",
      "labels": {
        "env": "prod",
        "job": "service1"
      },
      "attributes": {
        "BaseUrl": "https://service1.example.townsuite.com",
        "HealthCheck": "/healthz/ready",
        "PrometheusMetricsUrl": "/metrics",
        "DataCenter": "yyz1"
      }
    },
    {
      "name": "Service 1 — secondary",
      "id": "9c0a7b55-1d2e-4a88-b0a1-7c2f9e3d4b55",
      "labels": { "env": "prod", "job": "service1" },
      "attributes": {
        "BaseUrl": "https://service1b.example.townsuite.com",
        "HealthCheck": "/healthz/ready"
      }
    }
  ]
}
services[] one entry per instance

Each instance becomes one target group in the output file. An instance without a BaseUrl attribute is skipped.

labels passed straight through

Copied verbatim onto the Prometheus target. Use them for job, env and anything else you group or alert by.

Recognized attributes

BaseUrl all outputs
Scheme and host (plus optional base path) of the instance. Required — instances without it are skipped.
HealthCheck targets_v2.json
A health endpoint, appended to BaseUrl.
HealthCheck_<suffix> targets_v2.json
Additional health endpoints. Any suffix, any number of them — only the HealthCheck_ prefix is matched.
ExtraHealthCheck targets_v2.json
Path or absolute URL returning a JSON array of extra health endpoints for this instance.
ExtraHealthCheck_Prefix targets_v2.json
Optional path inserted between BaseUrl and each endpoint returned by ExtraHealthCheck.
PrometheusMetricsUrl targets_prometheus_metrics.json
Metrics path, appended to BaseUrl.
OpenTelemetryUrl targets_open_telemetry_metrics.json
OpenTelemetry metrics path, appended to BaseUrl. Declaring it makes the instance a scrape target; instances without it are left out of that file.
attributes is free-form — unrecognized keys such as DataCenter are ignored, so you can return whatever else your inventory tracks. Slashes are normalized when a path is joined to BaseUrl, so https://host/ + /healthz and https://host + healthz both give https://host/healthz.
Multiple endpoints

Several health endpoints for one site

Because attributes is a JSON object its keys must be unique — you cannot repeat HealthCheck. There are three ways to list more than one endpoint for a single instance, and they can be combined. All resulting URLs are collected into that instance's single targets array, sorted and de-duplicated.

A fixed list, spelled out in the inventory. Add any suffix after HealthCheck_ — the suffix is a label for you, the tool only matches the prefix and accepts any number of them.

discover response · services[0]
{
  "labels": { "Env": "prod", "Job": "service1" },
  "attributes": {
    "BaseUrl": "https://service1.example.townsuite.com",
    "HealthCheck": "/healthz/ready",
    "HealthCheck_live": "/healthz/live",
    "HealthCheck_database": "/healthz/database",
    "HealthCheck_queue": "/healthz/queue"
  }
}
targets_v2.json
[
  {
    "targets": [
      "https://service1.example.townsuite.com/healthz/database",
      "https://service1.example.townsuite.com/healthz/live",
      "https://service1.example.townsuite.com/healthz/queue",
      "https://service1.example.townsuite.com/healthz/ready"
    ],
    "labels": { "Env": "prod", "Job": "service1" }
  }
]

A list the site returns itself — use this when the site knows its own endpoints and you would rather not update the inventory every time one is added. The value is a path on the site or an absolute URL.

discover response · services[0].attributes
{
  "BaseUrl": "https://service1.example.townsuite.com",
  "HealthCheck": "/healthz/ready",
  "ExtraHealthCheck": "/healthz/extralookups"
}
GET https://service1.example.townsuite.com/healthz/extralookups
["hello/world", "world/hello"]
targets_v2.json
[
  {
    "targets": [
      "https://service1.example.townsuite.com/healthz/ready",
      "https://service1.example.townsuite.com/hello/world",
      "https://service1.example.townsuite.com/world/hello"
    ],
    "labels": { "Env": "prod", "Job": "service1" }
  }
]
The full contract for that lookup — and what happens when it fails — is in Extra health checks below.

When the returned array holds names rather than full paths, add a prefix that sits between the base URL and each returned endpoint.

discover response · services[0].attributes
{
  "BaseUrl": "https://service1.example.townsuite.com",
  "HealthCheck": "/healthz/ready",
  "ExtraHealthCheck": "/healthz/extralookups",
  "ExtraHealthCheck_Prefix": "/healthz/ready/extras"
}
targets_v2.json · same ["hello/world", "world/hello"] response
[
  {
    "targets": [
      "https://service1.example.townsuite.com/healthz/ready",
      "https://service1.example.townsuite.com/healthz/ready/extras/hello/world",
      "https://service1.example.townsuite.com/healthz/ready/extras/world/hello"
    ],
    "labels": { "Env": "prod", "Job": "service1" }
  }
]

Extra health checks — the site owns the list

ExtraHealthCheck inverts who owns the list. Rather than the inventory enumerating every endpoint, the site exposes one endpoint that returns the health endpoints it manages, and the tool expands it on every pass — add or remove one in the service and Prometheus follows on the next cycle, with no inventory change and no restart.

Declared by
ExtraHealthCheck in an instance's attributes — one value, not a list. One lookup per instance.
Request
GET, no body. Sends the SettingsV2 AuthHeader as Authorization and the configured UserAgent, under HttpTimeoutInSeconds.
Response
HTTP 2xx with a JSON array of strings. Nothing else is accepted — not an object, not an array of objects.
Each string
A path, joined to BaseUrl (or to BaseUrl + ExtraHealthCheck_Prefix). Leading slashes are optional.
the lookup endpoint, service side
app.MapGet("/healthz/extralookups", () => new[]
{
    "/healthz/database",
    "/healthz/queue",
    "/healthz/upstream/billing"
});

How one pass expands it

  1. The instance's HealthCheck and HealthCheck_* values are joined to BaseUrl and added first.
  2. The lookup URL is resolved — a value starting with http is used verbatim, anything else is joined to BaseUrl. A path asks the site itself what it manages; an absolute URL lets another service answer on its behalf.
  3. Each string in the returned array is joined to BaseUrl, or to BaseUrl + ExtraHealthCheck_Prefix when that attribute is set.
  4. Everything lands in that instance's single targets array, sorted, sharing its labels. Static endpoints are added first, so a discovered endpoint loses any duplicate or substring collision.

Failure behaviour

Lookup fails static endpoints survive

A non-success status, a transport failure, or a malformed body is logged, and the instance keeps every target already collected from HealthCheck / HealthCheck_*. Only the discovered extras are missing from that pass.

Nothing to fall back on previous file is kept

If the instance declared nothing but ExtraHealthCheck, a failed lookup loses the whole instance — so the pass is marked incomplete and the previous file is kept rather than published without it.

An absent() rule is still a useful backstop, since it catches a target disappearing for any reason — including a service legitimately removed from the inventory:

prometheus rules
- alert: HealthCheckTargetMissing
  expr: absent(probe_success{job="health_checks_service_discovery"})
  for: 15m

Cost and trust

Request volumeper cycle
The lookup runs once per instance, per cycle. Ten instances means ten requests every DelayInSeconds — there is no caching, even when they all point at the same absolute URL.
AuthHeaderis forwarded
The header configured for the inventory source is sent to the lookup URL too. If it is a shared credential, every host named by a BaseUrl — and any absolute URL you put in ExtraHealthCheck — receives it.
Returned pathshost-bound
Entries are appended to BaseUrl and never used as full URLs, so a compromised lookup cannot redirect probing at an unrelated host — but it can point the probe at any path on its own. IgnoreList suppresses specific paths.

Things to watch

One host per instance

Every HealthCheck_* value is appended to that instance's BaseUrl. Endpoints on a different host belong to another services[] entry with its own BaseUrl.

Avoid nested paths

Duplicates are dropped, and so is any URL already contained in a target in the same group. Prefer distinct, fully-specified paths over paths that nest inside one another.

One instance always yields one target group sharing that instance's labels. Prometheus probes each URL in targets separately, so with the relabel rules in the Prometheus section the instance label holds the full probed URL — that is what keeps /healthz/ready and /healthz/database apart in alerts. A single noisy endpoint can be dropped with IgnoreList without removing the whole instance.
Failure handling

Temporary failures do not empty the target files

A lost lookup means the targets are unaccounted for, not that there are none. Publishing an emptier file would be worse than publishing nothing: prometheus drops the targets it no longer sees, stops probing them, and an alert on a missing target does not fire — it just goes quiet. So every pass reports whether it was complete, and an incomplete pass does not overwrite the previous file.

What gets written

Complete
The new list replaces the file, as always.
Incomplete, previous file exists
Nothing. The previous file stays in place, a warning is logged, and the next cycle tries again.
Incomplete, no previous file
Whatever was discovered — holding back on a first run would leave prometheus with no file at all.
Incomplete past MaxStaleCycles
The incomplete list is written anyway. Only applies when MaxStaleCycles is greater than zero.

What makes a pass incomplete

ServiceListUrl fails
Non-success status or unparseable body — every target from that source is unaccounted for.
ServiceDiscoverUrl fails
For any service key — that service's instances are unaccounted for.
ExtraHealthCheck fails with nothing to fall back on
The instance declared no static HealthCheck, so it is gone from this pass entirely.
An unusable source entry
A null entry, or one with a blank ServiceListUrl.
A transport-level exception (timeout, DNS, TLS) is covered by the same guarantee from the other direction — it unwinds before the temp file is moved into place, so the previous file survives untouched while the outer handler logs Crashing and exits for the service manager to restart.
the log line to watch for
warn: targets_v2.json pass was incomplete, keeping the previous file
      so the endpoints stay in prometheus (incomplete pass 3)
Default MaxStaleCycles: 0

Holds the last good files for as long as the lookups keep failing, because going blind is the worse failure. The trade-off: a service genuinely removed from the inventory keeps being probed while the source is broken.

Converging MaxStaleCycles > 0

Accepts the incomplete list after that many consecutive failures. 48 with a 30-minute DelayInSeconds gives up after a day of continuous failure.

Deploy

Run as a service

Keep it running in the background and restart it automatically. Use NSSM on Windows or a systemd unit on Linux.

  1. Copy the executable to C:\prometheus\TownSuite.Prometheus.FileSdConfigs
  2. Edit appsettings.json
  3. Register it with NSSM
PowerShell
nssm install TownSuite.Prometheus.FileSdConfigs C:\prometheus\...\TownSuite.Prometheus.FileSdConfigs.exe
nssm set TownSuite.Prometheus.FileSdConfigs AppDirectory C:\prometheus\TownSuite.Prometheus.FileSdConfigs
net start TownSuite.Prometheus.FileSdConfigs
  1. Extract the release to /opt/TownSuite.Prometheus.FileSdConfigs
  2. Create the unit file below at /etc/systemd/system/
  3. Enable & start with systemctl
townsuite-prometheus-filesdconfigs.service
[Unit]
Description=TownSuite Prometheus FileSdConfigs

[Service]
ExecStart=/opt/TownSuite.Prometheus.FileSdConfigs/TownSuite.Prometheus.FileSdConfigs
Restart=always
RestartSec=10
StandardOutput=syslog
StandardError=syslog
SyslogIdentifier=TownSuite.Prometheus.FileSdConfigs
WorkingDirectory=/opt/TownSuite.Prometheus.FileSdConfigs/

[Install]
WantedBy=multi-user.target
bash
systemctl enable townsuite-prometheus-filesdconfigs.service
systemctl start townsuite-prometheus-filesdconfigs.service
Integrate

Wire it into Prometheus

Add file_sd_configs jobs to prometheus.yml — one scrapes the discovered metrics endpoints, the other runs health checks through the Blackbox exporter.

metrics_service_discovery
Reads targets_prometheus_metrics.json and relabels scheme, address & metrics path from each URL.
health_checks_service_discovery
Reads targets_v2.json and probes each target via the Blackbox exporter at :9115.
open_telemetry_service_discovery
Reads targets_open_telemetry_metrics.json with the same relabelling — a plain scrape, so an instance appears only if it declares OpenTelemetryUrl.
/etc/prometheus/prometheus.yml
- job_name: metrics_service_discovery
  scheme: http
  file_sd_configs:
  - files:
     - /opt/TownSuite.Prometheus.FileSdConfigs/targets_prometheus_metrics.json
  scrape_interval: 60s
  relabel_configs:
    - source_labels: [__address__]
      regex: '^(https?)://([^/]+)(/.*)?'
      target_label: __scheme__
      replacement: '${1}'
    # ...address & metrics_path relabels follow

# OpenTelemetry endpoints — same relabelling, different file
- job_name: open_telemetry_service_discovery
  scheme: http
  file_sd_configs:
  - files:
     - /opt/TownSuite.Prometheus.FileSdConfigs/targets_open_telemetry_metrics.json
  scrape_interval: 60s
  relabel_configs:
    # same scheme / address / metrics_path relabels as above

# Health checks via the Blackbox exporter
- job_name: health_checks_service_discovery
  file_sd_configs:
  - files:
     - /opt/TownSuite.Prometheus.FileSdConfigs/targets_v2.json
  metrics_path: /probe
  params:
    module: [http_2xx]
  relabel_configs:
    - source_labels: [__address__]
      target_label: __param_target
    - target_label: __address__
      replacement: 127.0.0.1:9115
The OpenTelemetry job is a plain scrape, not a blackbox probe. Point it at an OTel Collector's prometheus receiver instead of Prometheus itself if the collector owns the scrape.