CraftFileGate
CraftFileGate is a file server: an SFTP client, an HTTP script or a browser
puts files on it and takes files from it, and the server stores them on one
or more storages. It replaces a classic SFTP server (ProFTPD, OpenSSH
internal-sftp) when the files do not all live on a local disk, or when the
accounts come from an identity provider.
What it does
| Need | Answer |
|---|---|
| Put and take files | SFTP, REST file API, web file explorer |
| Store the files | local disk, upstream SFTP server, S3 and compatibles, HDFS via Knox (read only) |
| Authenticate | password, SSH public key, JWT token (secret, public key or JWKS) |
| Grant rights | roles, which mount storages and grant rights on them by path; refusal by default |
| Monitor | audit trail, JSON logs, Prometheus metrics, OpenTelemetry traces |
| Administer | admin API and web console: sessions, bans, configuration in force |
| Deploy | a static binary, Docker images, a Helm chart |
The words of this guide
| Word | Meaning |
|---|---|
| backend | a named storage, described in [[backends]] |
| role | a set of mounts; a user receives one or more |
| mount | a backend seen by the user at a path (mount_path), starting from a subdirectory (home_dir), with its ACL |
| ACL | the rights (read, write, list, delete, rename) by path, within a mount |
| door | a way in: SFTP, REST API, file explorer, admin console |
The chapter Users, roles and mounts explains them with examples.
Quick start
A server on a local disk, a user alice who sees only her directory.
1. Generate a host key
ssh-keygen -t ed25519 -f /etc/craft-file-gate/host_ed25519 -N ""
A key listed in host_keys but missing refuses startup.
2. Write the configuration
Create /etc/craft-file-gate/config.toml:
#:schema ./config.schema.json
[server]
shutdown_grace_period_secs = 30
# max_sessions_per_user = 10 # optional, unlimited by default
[sftp]
listen = "0.0.0.0:2222"
host_keys = ["/etc/craft-file-gate/host_ed25519"]
# login_grace_secs = 120 # unauthenticated connection cut after this (0 = never)
[auth]
jwt_sentinel_username = "jwt"
timeout_secs = 5
# roles_file = "/etc/craft-file-gate/roles.toml" # optional, otherwise inline [[roles]]
# authz_base_url = "https://authz.internal" # optional, remote mapping service
[auth.jwt]
secret = "changez-moi-en-production"
# public_key_file = "/etc/craft-file-gate/jwt_public.pem" # alternative to the secret
# jwks_url = "https://idp.interne/.well-known/jwks.json" # alternative: identity provider keys
# jwks_refresh_interval_secs = 3600 # default, minimum 1
# username_path = "/sub" # default
# authorities_path = "/groups" # default
[auth.methods]
jwt = { enabled = true }
# Password and public key are proofs of the local store, not separate
# methods: you declare them under `local`.
local = { enabled = true, password = true, pubkey = true }
[admin]
listen = "127.0.0.1:8080"
bearer_token = "changez-moi-en-production"
[log]
level = "info"
format = "json"
# ── Storage and rights ──
[[backends]]
name = "local"
type = "local"
root = "/srv/sftp"
[[roles]]
name = "utilisateurs"
[[roles.mounts]]
backend = "local"
home_dir = "/{username}" # alice sees /srv/sftp/alice as her root /
create_home = true
[[roles.mounts.acl]]
path = "/"
rights = ["read", "write", "list", "delete", "rename"]
recursive = true
# ── Local users ──
[[users]]
username = "alice"
# password: changez-moi-en-production
password_hash = "$argon2id$v=19$m=19456,t=2,p=1$QzIrOEdtBZI4UL4ddXDbkg$yudDYsZSLamFZHWeKlFpFWtkP4kX6ALV8GNKs3pEXBk"
# authorized_keys = ["ssh-ed25519 AAAA... alice@poste"]
authorities = ["utilisateurs"]
The JWT secret, the admin token and alice’s password (changez-moi-en-production) are public. The server starts with them, but emits one WARN per example value still in place. To replace alice’s [[users]] block:
echo -n "mot-de-passe" | craft-file-gate hash-password --user alice
The process must be able to write to /srv/sftp: create_home creates alice/ there at her first login.
3. Start the server
craft-file-gate --config /etc/craft-file-gate/config.toml
For Docker and Kubernetes, see the guide (“Operating”).
4. Connect
sftp -P 2222 alice@votre-serveur
ls shows the content of /srv/sftp/alice, and nothing above it.
What next
| To | Read |
|---|---|
| give everyone what they should see | Users, roles and mounts, Recipes |
| connect S3 or an upstream SFTP server | Backends |
| connect an identity provider | Authentication |
| deploy in a container | Docker, Kubernetes |
| understand a refusal | Troubleshooting, Audit |
| look up a key | Configuration reference |
Users, roles and mounts
user ──authorities──▶ role ──▶ mount ──▶ backend (storage)
│
└── ACL (rights per path)
No right is implicit: a backend that no role mounts is visible to nobody, and the admin token gives access to no file.
A complete example
[[backends]] # one storage, described once
name = "disque"
type = "local"
root = "/srv/sftp"
[[roles]]
name = "utilisateurs"
[[roles.mounts]] # at least one per role
backend = "disque" # a [[backends]] name
home_dir = "/{username}" # where the mount starts on the storage
create_home = true
[[roles.mounts.acl]] # the rights, relative to the mount
path = "/"
rights = ["read", "write", "list", "delete", "rename"]
recursive = true
[[users]]
username = "alice"
password_hash = "$argon2id$..." # craft-file-gate hash-password
authorities = ["utilisateurs"] # alice's roles
alice sees /; this / is /srv/sftp/alice on the disk.
The keys of a role and of a mount
All the keys, with their hot reload: Reference [[roles]],
[[roles.mounts]].
Where the files go
A client path goes to the mount whose mount_path is its longest prefix;
the rest is appended to home_dir, on the storage
(what home_dir points to on each type).
.. never climbs above home_dir.
{username}: one directory per user
home_dir = "/{username}" gives each user their own directory, with a
single role. The name, never sanitized, is made of A-Z a-z 0-9 . _ -, 1 to
64 characters, with no leading dot.
One mount or several
| Mounts | What the user sees |
|---|---|
one, at / | the backend tree, starting from home_dir |
several, under names (/disque, /archives) | / lists the mounts; each one shows its backend |
nested (/partenaires/acme) | / and /partenaires are intermediate directories |
The paths between mounts are synthetic directories. A rename stays within its mount.
Combining roles
A user often has several roles. Their session has the mounts of all of them.
| The mounts of two roles | In the session |
|---|---|
mount_path values none of which is a prefix of another (/a, /b) | they coexist |
identical: same mount_path, backend, home_dir, and same max_file_mb, create_home, hidden_stores | a single mount, ACLs merged |
Two different mounts of the same backend have disjoint home_dir values
(/a and /b, not / and /a): each file is reached under one ACL only.
The doctrine: a role at / is complete
- a role at
/is complete: on its own it gives everything its user sees, and it combines with no other; - a role meant to be combined mounts under a name (
/archives).
See the recipes.
Where roles come from
| Source | How | The roles |
|---|---|---|
| local user | authorities of their [[users]] entry | names of [[roles]] |
| JWT token | the authorities_path claim (default /groups) | names of [[roles]], or authorities to translate |
| authorization service | authz_base_url translates the unknown authorities | names of [[roles]] |
Roles are always defined in [[roles]] or roles_file. Only the
authorities that are not role names go to the service.
The authorization service contract
POST <authz_base_url>/authz/resolve
Content-Type: application/json
{"authorities": ["CN=ACME-Partners,OU=Groups"]}
Expected response: 200 and
{"roles": ["partenaire-acme"]}
authz_base_url and timeout_secs: Reference [auth].
The named roles are added to those found locally; a name that [[roles]]
does not define is ignored, and a service that is down leaves the local roles.
An SFTP session builds all its backends at authentication and keeps its mounts until it ends, even after a hot reload; a REST request builds the backend of the mount it reaches.
What you will see
| When | Line |
|---|---|
startup, role at {username} | INFO role gives each user their own home directory |
| startup, role that mounts several backends | INFO role mounts several backends: its users hold the credentials of all of them ..., fields role, backends, writable |
create_home creates a directory | INFO created missing home directory |
A role that mounts several backends holds their credentials: keep it read only, for a dedicated user authenticated by key.
Mount recipes
Each recipe assumes this backend:
[[backends]]
name = "disque"
type = "local"
root = "/srv/sftp"
1. One directory per user
[[roles]]
name = "utilisateurs"
[[roles.mounts]]
backend = "disque"
home_dir = "/{username}"
create_home = true
acl = [{ path = "/", rights = ["read", "write", "list", "delete", "rename"], recursive = true }]
alice sees / = /srv/sftp/alice, created at her first login. A role
at / does not combine.
2. A partner: inbound drop, outbound pickup
The partner drops into /in without being able to read back, and picks up
from /out.
[[roles]]
name = "partenaire-acme"
[[roles.mounts]]
backend = "disque"
home_dir = "/partenaires/acme"
max_file_mb = 500
hidden_stores = { enabled = true, prefix = ".in.", extension = "" }
acl = [
{ path = "/in", rights = ["write", "list"], recursive = true },
{ path = "/out", rights = ["read", "list", "delete"], recursive = true },
]
/in is /srv/sftp/partenaires/acme/in; an upload there is published at
the end of the transfer.
For ls / to show in and out, add a non-recursive entry:
path = "/", rights = ["list"].
3. The operations account: every storage, read only
An account that sees each backend at the root and cannot modify anything.
[[backends]]
name = "archives"
type = "s3"
bucket = "archives"
region = "eu-west-3"
prefix = "sftp/"
credentials = { type = "iam_role" }
[[backends]]
name = "acme-amont"
type = "sftp"
host = "sftp.acme.example"
host_key_fingerprint = "SHA256:2hZbXq1b5Xb3vN2mQ6w7l3Fv0tXk4yJ8aUeYp9rS0cE" # ssh-keygen -lf, checked out of band
auth = { type = "password", username = "relais", password = "changez-moi" }
[[roles]]
name = "exploitant"
[[roles.mounts]]
backend = "disque"
mount_path = "/disque"
acl = [{ path = "/", rights = ["read", "list"], recursive = true }]
[[roles.mounts]]
backend = "archives"
mount_path = "/archives"
acl = [{ path = "/", rights = ["read", "list"], recursive = true }]
[[roles.mounts]]
backend = "acme-amont"
mount_path = "/acme-amont"
home_dir = "/depot"
acl = [{ path = "/", rights = ["read", "list"], recursive = true }]
[[users]]
username = "exploitation"
password_hash = "$argon2id$..." # of a random, discarded password: only the key is used
authorized_keys = ["ssh-ed25519 AAAA... exploitation@poste"]
authorities = ["exploitant"]
ls / shows acme-amont archives disque; /archives/2024/x is the S3 key
sftp/2024/x.
At startup, INFO role mounts several backends: ... gives role=exploitant
and writable=[]: check it after each change to the role.
4. Roles that combine
[[roles]]
name = "equipe-a"
[[roles.mounts]]
backend = "disque"
mount_path = "/equipe-a"
home_dir = "/equipes/a"
acl = [{ path = "/", rights = ["read", "list"], recursive = true }]
[[roles]]
name = "equipe-b"
[[roles.mounts]]
backend = "disque"
mount_path = "/equipe-b"
home_dir = "/equipes/b"
acl = [{ path = "/", rights = ["read", "write", "list"], recursive = true }]
A token whose groups claim is ["equipe-a", "equipe-b"] opens a session
with /equipe-a (read) and /equipe-b (read and write).
Two roles on the same mount merge their ACLs: acme-depot (/in) and
acme-releve (/out) on the mount of case 2 give /in and /out.
5. Organized mounts: /partenaires/<name>
With acme-amont from case 3:
[[roles]]
name = "gestion-partenaires"
[[roles.mounts]]
backend = "disque"
mount_path = "/partenaires/acme"
home_dir = "/partenaires/acme"
acl = [{ path = "/", rights = ["read", "list"], recursive = true }] # "/": the root of the mount
[[roles.mounts]]
backend = "acme-amont"
mount_path = "/partenaires/amont"
home_dir = "/depot"
acl = [{ path = "/", rights = ["read", "list"], recursive = true }]
/ and /partenaires are synthetic; /partenaires/amont is the
upstream’s /depot.
ACL
The ACL of a mount grants rights, path by path; everything else is refused.
It is written under [[roles.mounts.acl]], and its path is
relative to the mount.
Writing an ACL
[[roles.mounts]]
backend = "disque"
mount_path = "/partenaire"
home_dir = "/partenaires/acme"
[[roles.mounts.acl]]
path = "/" # /partenaire for the user
rights = ["list"] # not recursive: this folder only
[[roles.mounts.acl]]
path = "/in" # /partenaire/in
rights = ["write", "list"]
recursive = true # and everything below it
All the keys: Reference [[roles.mounts.acl]].
An ACL governs only its own mount, even on a backend mounted twice.
The rights
Each operation requires its right on each path it names:
| Operation | Required right |
|---|---|
upload, mkdir | write on the path, and on each parent that must be created |
| download | read |
| listing | list on the listed directory |
stat (SFTP), HEAD (REST) | read or list |
| deleting a file, an empty directory | delete |
| deleting a non-empty directory | delete on the whole tree |
| rename | rename on the source and the destination, and write where the moved content lands |
A listing shows the whole directory, without filtering by the ACL. list on
/in does not give list on /.
Which entry decides
- Refusal by default: a path that no entry governs is refused.
- The most specific entry decides: the one for the exact path, otherwise
the closest
recursiveentry above it. - A more specific entry that does not grant the right refuses, even under a broader entry that grants it.
[[roles.mounts.acl]]
path = "/"
rights = ["read", "write", "list", "delete", "rename"]
recursive = true
[[roles.mounts.acl]]
path = "/archives"
rights = ["read", "list"]
recursive = true
Here everything is writable, except /archives and its content, read only.
A move is judged again on what it carries: a folder that contains a protected subfolder does not move to a place governed by a broader entry, where the subtree would lose its protection.
Several roles on one mount
Two roles on the same mount put their ACLs together:
- two entries at the same path merge their rights, and a single
recursivemakes the union recursive; - at different paths, the most specific entry decides, whichever role it comes from.
A narrow entry from one role can therefore restrict what a broad entry from
another granted. The result is shown in the console, in the roles tab of a
session (GET /admin/sessions/<id>/roles).
Synthetic directories
With several mounts under names, a path under no mount (/,
/partenaires) is synthetic: it belongs to no backend.
| Operation | On a synthetic path |
|---|---|
| listing | the mounts and the synthetic directories it contains |
stat | a directory, dated from the start of the session |
?rights (REST) | ["list"] |
| any other operation | refused |
path under which there is no mount (/inconnu) | not found |
A mount point (/partenaires/acme) is listed according to its ACL, but is
not written, deleted or renamed.
Name case
When a backend’s ACL folds case
(case_insensitive),
/Archives and /archives are a single path: write it once per mount.
Doors: SFTP, REST API, file explorer, admin console
The file doors (SFTP, REST API, file explorer) share users, roles, mounts and ACLs: a right granted is granted everywhere. The admin console opens no file. SFTP holds sessions; the REST API has none, and the console shows its long transfers in progress (its activity).
| Door | For whom | Listens on | Credential | Section |
|---|---|---|---|---|
| SFTP | SFTP clients (OpenSSH sftp, WinSCP, FileZilla, rclone), sshfs | [sftp] listen | password, SSH key, JWT (sentinel name) | [sftp] |
| REST file API | HTTP scripts, applications | [admin] listen, under [api] prefix | Basic, Bearer (JWT), ticket | [api] |
| Web file explorer | users in a browser | [admin] listen, at [api.ui] path | that of the API, kept in memory | [api.ui] |
| Admin console and API | operators | [admin] listen, or [admin] control_listen | local account (password) or JWT with a role from [[admin.roles]], fallback bearer_token | [admin] |
SFTP
[sftp]
listen = "0.0.0.0:2222"
host_keys = ["/etc/craft-file-gate/host_ed25519"]
- Only the
sftpsubsystem is served: no shell, noscp, noexec, no port or agent forwarding. - With JWT: the sentinel name (
jwtby default) and the token as password,sftp -P 2222 jwt@serveur. - The keys: The SFTP door.
REST file API
[admin]
listen = "0.0.0.0:8080"
bearer_token = "changez-moi-en-production"
[api]
enabled = true
prefix = "/api/v1/files"
# list a directory
curl -u alice:mot-de-passe 'http://serveur:8080/api/v1/files/in?list'
# upload a file
curl -u alice:mot-de-passe -T rapport.csv http://serveur:8080/api/v1/files/in/rapport.csv
# download
curl -u alice:mot-de-passe -o rapport.csv http://serveur:8080/api/v1/files/in/rapport.csv
| Request | Effect |
|---|---|
GET <path>?list&offset=&limit= | paginated listing; limit 100 by default, 10000 at most |
GET <path> | download; a single-interval Range resumes it (206) |
HEAD <path> | metadata |
PUT <path> | upload of the whole file |
PUT <path>?mkdir | directory creation |
DELETE <path> | deletion |
POST <path>?rename=<destination> | rename |
GET <path>?rights | {"path", "rights", "home"}: the caller’s rights on this path, without touching the storage |
POST <file>?ticket | 201 {"url": "<prefix>/<file>?ticket=<token>", "expires_in": 600}; read right; 32 live tickets per user, and a ticket that still has 60 s left is returned again |
GET <file>?ticket=<token> | the file, without Authorization, as many times as wanted for 600 s (less if the JWT of the request expires before), the ACL judged again each time |
All the keys: Reference [api].
- One credential per request; no cookie, no session.
- A downloaded file goes out as
attachment,Cache-Control: no-store,Content-Security-Policy: sandbox,Accept-Ranges: bytes. A leaked ticket URL downloads that file until it expires.
Web file explorer
[api.ui]
enabled = true
path = "/files"
A page at http://serveur:8080/files, a client of the API: same
credentials, same ACLs. See File explorer.
Admin console and API
[admin]
listen = "127.0.0.1:8080"
bearer_token = "changez-moi-en-production"
# control_listen = "127.0.0.1:8081" # a separate door for administration
The console is at / of the admin listener: sessions and mounts, bans,
configuration without secrets, logs, metrics, revocations. You log in with
a username and password, or with a token. /metrics and the probes are on
the same listener, or on control_listen, or on [server] probes_listen:
see One door or two.
What you will see
| Line | Meaning |
|---|---|
INFO SFTP server listening | the SFTP door is open |
INFO starting admin API server (HTTP) (or (HTTPS)) | the admin listener is open |
INFO every configured door started | every configured door is serving (field doors) |
audit connection_accepted | a successful authentication, on any door (on each request for REST) |
Backends
A backend is a named storage, described once in [[backends]]; mounts
refer to it by name.
[[backends]]
name = "disque" # the name that mounts refer to
type = "local" # local, sftp, s3 or webhdfs
root = "/srv/sftp" # the type-specific keys
[[roles]]
name = "utilisateurs"
[[roles.mounts]]
backend = "disque"
home_dir = "/{username}"
acl = [{ path = "/", rights = ["read", "write", "list"], recursive = true }]
[[backends]] are written in config.toml or in roles_file, where they
are reloaded at runtime with the roles.
The types
| Type | type = | The storage | What home_dir points to | Writes |
|---|---|---|---|---|
| Local | "local" | a directory on the server | a subdirectory of root | yes |
| SFTP proxy | "sftp" | an upstream SFTP server, under a service account | a directory on the upstream | yes |
| S3 and compatibles | "s3" | an AWS S3 bucket, MinIO, Garage, Scaleway… | a key prefix, after prefix | yes |
| WebHDFS (Knox) | "webhdfs" | HDFS through Apache Knox | an HDFS directory | read-only |
The published binaries and images carry all four types (The binary’s features).
The keys of every backend
All keys: Reference [[backends]].
A subtable ([backends.auth], [backends.credentials]) attaches to the
[[backends]] before it: write it right after it. GET /admin/config
and the Configuration tab show each backend, secrets replaced by ***.
Keys common to all types
What they do on each type is on the type’s page; their default and their effect, in the same Reference.
hidden_stores resolves from the mount, then the backend, then [server.hidden_stores];
stale_partials from the backend, then [uploads.stale_partials]; key by key
(Uploads).
Reservation across instances
The first upload to a destination holds it until it ends
(why). Across instances, this is a
lock <lock_prefix><name> placed next to the destination. lock_prefix:
- is 8 characters to 64 bytes long, with no
/and no control character; - also names the server’s throwaway names (
PPnew.<token>,PPprobe.,PPstale.for a prefixP); - reserves for the server any name that starts with it, regardless of case,
like
.craftfilegate-upload...: choose one that no user file carries.
Availability
Each instance visits each backend in the background ([server.backend_probe]): up, down after failures_before_down failed visits, unknown before the first one; a WARN at each change (Troubleshooting), the state in GET /admin/health, the Instances view and craftfilegate_backend_up (Metrics). On S3, a visit is a billed ListObjectsV2, every 15 s per instance by default.
Local
The local backend serves a directory of the server’s file system, its
root (root). Each mount has its own subdirectory there, home_dir.
[[backends]]
name = "disque"
type = "local"
root = "/srv/sftp"
[[roles]]
name = "utilisateurs"
[[roles.mounts]]
backend = "disque"
home_dir = "/{username}"
create_home = true
acl = [{ path = "/", rights = ["read", "write", "list", "delete", "rename"], recursive = true }]
alice uploads /rapport.txt: the file is /srv/sftp/alice/rapport.txt.
The keys
All keys: Reference [[backends]] type = “local”;
common keys: Backends.
The root and symbolic links
No operation leaves the mount’s root + home_dir, whatever links are
there, including what the server does itself (in-flight file, lock, sweep).
A link to another mount’s home_dir leaves it just as much as a link
outside the root.
| Setting | A link within the mount | A link that leaves it |
|---|---|---|
follow_symlinks = false (default) | never followed | never followed |
follow_symlinks = true | followed if it stays under root + home_dir | never followed |
- A link is listed as a link (type
l;"is_symlink": truein REST), or, when followed, as its target. Deleting or renaming acts on the link. - With
hidden_stores, an upload onto a file link replaces the link; the target stays intact. - No door creates links: they come from another process or a restore.
- The root and its ancestors are trusted (a root through a link is followed): writable by the administrator only.
- On a shared volume (NFS…), mount the root
nosymfollow(Linux 5.10 and later).
Case of ACL paths
On a storage that ignores case (NTFS, APFS and HFS+ by default, SMB/CIFS,
ext4 or tmpfs casefold), the ACL compares folded paths: NFD, removal
of ignorable code points (U+200B…), full case folding
(straße = strasse), NFC.
| Setting | Comparison of ACL paths |
|---|---|
case_insensitive = true | folded |
case_insensitive = false | byte by byte, without normalization |
| key absent | a probe decides, at startup and at each roles reload |
The probe creates a lowercase file and looks for it in uppercase in
each directory where an entry resolves (folding is set per directory,
chattr +F): one folding directory makes the backend fold, an
inconclusive result counts as folding. A missing directory or one under
{username} is not probed, and a case-sensitive APFS volume stays
insensitive to normalization: when in doubt, case_insensitive = true.
Where the ACL folds, and on a Windows server, a short 8.3 name (PROTEG~1)
or one ending in . or a space is not a valid path: disable short names
(fsutil 8dot3name).
The common keys on a local backend
| Key | On a local backend |
|---|---|
home_dir (mount) | a subdirectory of root: /partenaire/in/x of a mount at /partenaire with home_dir = "/partenaires/acme" is /srv/sftp/partenaires/acme/in/x |
create_home (mount) | true: creates home_dir if missing, when the session opens, one level at a time, never through a link |
hidden_stores | a file being transferred next to the destination, published by rename |
stale_partials | the sweep of those files, on the file system’s clock |
cross_instance_reservation | a lock file next to the destination; false by default on Windows: one instance per root |
lock_prefix | the name of those locks |
The process acts under its own uid; file modes follow its umask.
On Linux 5.6 and later, the kernel enforces the confinement (openat2(2),
RESOLVE_BENEATH); without openat2, each component is opened
O_NOFOLLOW (reported once as INFO); on Windows, the parent is checked
just before the call.
Performance and limits
- A rename and the publication of an upload use
RENAME_NOREPLACE. Where the file system lacks it (NFS, FUSE including sshfs, 9p), they fall back torename(2): a concurrent overwrite is not detected there, and the audit of an upload that overwrites saysreplaced=unknown. - No
fsync: a success does not promise durability on disk. mkdircreates one level; the missing parents of an upload are created one by one, each judged by the ACL.- Deleting a tree is not counted: the audit says
removed=unknown.
What you will see
| When | Line |
|---|---|
| startup, case probe | INFO with backend, acl_paths = folded or exact |
create_home creates a directory | INFO created missing home directory, field home_dir |
| startup, sweep of leftovers | INFO stale in-flight files: ..., fields backend, grace_secs, age_check |
SFTP proxy
The sftp backend stores files on an upstream SFTP server, reached over
SSH with a service account that users do not know.
[[backends]]
name = "acme-amont"
type = "sftp"
host = "sftp.acme.example"
host_key_fingerprint = "SHA256:2hZbXq1b5Xb3vN2mQ6w7l3Fv0tXk4yJ8aUeYp9rS0cE" # checked out of band
[backends.auth]
type = "password"
username = "relais"
password = "changez-moi"
[[roles]]
name = "acme"
[[roles.mounts]]
backend = "acme-amont"
home_dir = "/depot"
acl = [{ path = "/", rights = ["read", "write", "list"], recursive = true }]
A client of the acme role uploads /f.txt: the file is /depot/f.txt on
the upstream.
The keys
All keys: Reference [[backends]] type = “sftp”;
common keys: Backends.
Key authentication:
[backends.auth]
type = "private_key"
username = "relais"
private_key_pem = """
-----BEGIN OPENSSH PRIVATE KEY-----
...
-----END OPENSSH PRIVATE KEY-----
"""
Prefer an ed25519 or ECDSA key. An RSA key signs with rsa-sha2-512 or
rsa-sha2-256 depending on server-sig-algs, ssh-rsa (SHA-1) as a last resort.
The upstream host key
The proxy checks the upstream host key before authenticating: otherwise, a man in the middle would receive the service account credentials. The fingerprint:
ssh-keyscan -p 22 sftp.acme.example | ssh-keygen -lf -
- Pin the ED25519 fingerprint, failing that ECDSA, failing that RSA: this is the order in which the proxy negotiates.
- Verify it out of band: on the upstream,
ssh-keygen -lf /etc/ssh/ssh_host_ed25519_key.pub. - Lowercase
sha256:and trailing=padding are accepted. - A key presented as a certificate is judged on the key it certifies.
- Rotation: one fingerprint per backend; change it at switchover, the roles reload applies it to the next connections.
Uploads
With hidden_stores.enabled = true, an upload writes to an in-transfer file
on the upstream, next to the destination, then publishes it:
| Step | Request to the upstream | replaced |
|---|---|---|
| 1 | SSH_FXP_RENAME of the in-transfer file onto the destination; succeeds if it is free | no |
| 2 | destination taken: posix-rename@openssh.com, atomic, on a second SFTP channel opened at the first overwrite | yes |
| 3 | without this extension or a second channel: deletion of the destination, then rename | yes |
At step 3, the destination is missing for one round trip. Without
hidden_stores, the upload writes to the destination, and resume and append
are served (Uploads).
The service account can create, rename and delete in the destination directory. An overwrite replaces the inode (mode, owner, ACL, extended attributes lost) and takes twice the space until publication.
Common keys on an SFTP proxy
| Key | On an SFTP proxy |
|---|---|
home_dir (mount) | an absolute path on the upstream |
create_home (mount) | no effect: home_dir exists on the upstream |
hidden_stores | an in-transfer file on the upstream, published by rename |
stale_partials | the sweep of these files on the upstream, judged on the upstream clock through a probe file; at startup, it waits 30 s at most for the connection |
cross_instance_reservation | a lock file on the upstream, next to the destination; true by default, Windows included |
lock_prefix | the name of these locks; another prefix, without a leading dot, suits an upstream that refuses the default name |
case_insensitive | false by default: the upstream file system is not visible from here; a Windows or macOS upstream wants true |
Performance and limits
- Read: one
READrequest in flight per handle; a read that follows the previous one is served from the block already read. A download in order reads the file about once on the upstream. - Write: packets as large as the upstream accepts, eight in flight.
- One SSH connection per SFTP session and per REST request: the handshake weighs on short REST requests.
- The proxy speaks SFTP v3 (statuses 0 to 8), like OpenSSH.
- Upstream links are not told apart:
statfollows a link, a listing marks no entry as a link.
What you will see
| When | Line |
|---|---|
| session opening | INFO SFTP proxy: connecting, fields host, port |
| host key verified | INFO SFTP proxy: host key verified, fields upstream, fingerprint |
| session ready | INFO SFTP proxy: connected and ready, fields host, home_dir |
S3 and compatibles
The s3 backend stores files as objects in an AWS S3 bucket or in a
compatible storage (MinIO, Garage, Scaleway…).
[[backends]]
name = "archives"
type = "s3"
bucket = "archives"
region = "eu-west-3"
prefix = "sftp/"
endpoint_url = "https://minio.interne.example:9000" # without this key: AWS
[backends.credentials]
type = "static"
access_key_id = "AKIAEXEMPLE"
secret_access_key = "changez-moi"
[[roles]]
name = "archivistes"
[[roles.mounts]]
backend = "archives"
home_dir = "/{username}"
acl = [{ path = "/", rights = ["read", "write", "list"], recursive = true }]
alice uploads /in/x.txt: the object is the key sftp/alice/in/x.txt.
The keys
All keys: Reference [[backends]] type = “s3”.
iam_role: AWS_* variables, profile, instance or pod role. Common
keys: Backends.
Files and directories
| What the client sees | On the storage |
|---|---|
| a file | a key |
| a directory | a marker <key>/, or any key under <key>/ |
a mkdir | writes the marker |
| missing parents of an upload | one marker per level, each judged by the ACL like a mkdir |
| a directory without a marker | size 0, no date |
A key and a directory with the same name can coexist, except with
refuse_upload_over_directory = true.
| Operation | S3 requests |
|---|---|
| SFTP read | one HeadObject at open, then one GetObject with Range per read |
| REST download | one GetObject read as a stream |
| upload under 8 MiB | one PutObject with If-None-Match: *, at the end |
| upload of 8 MiB or more | a multipart upload, one 8 MiB part at a time |
| file rename | CopyObject with If-None-Match: *, then DeleteObject of the source |
| listing | ListObjectsV2 with the / delimiter, page after page |
| tree deletion | ListObjectsV2 by 1,000 keys, then DeleteObjects in batches |
An object is written whole: no resume (reput), no append. Only a file
can be renamed.
Bucket permissions
On the bucket and <prefix>/*:
| Action | For |
|---|---|
s3:ListBucket | list, tell a directory apart, find missing parents, mkdir, rename, delete a tree |
s3:GetObject | download, stat, and the check that precedes a deletion or a rename |
s3:PutObject | upload, mkdir, rename, the conditional write self-test |
s3:DeleteObject | delete, rename, clean up the self-test key |
s3:AbortMultipartUpload | abort a multipart upload that will not be completed |
s3:ListBucketMultipartUploads | sweep abandoned multipart uploads |
s3:ListMultipartUploadParts | date the last part of a sweep candidate |
A read-only backend only needs s3:ListBucket and s3:GetObject.
Multipart uploads
A multipart upload that is never completed is billed until it is aborted:
| When | Abort |
|---|---|
| an upload ends without completing | immediately, whatever the cause |
| clean shutdown | those still in flight, 5 s at most |
| a killed process | the sweep under <prefix>/: at startup, every grace_secs and at each upload, when the last part is older than grace_secs on the storage clock |
- Give each backend a
prefix: the sweep only works under it. - A slow upload sends an empty part every 15 s, which dates it for the
other instances; they have the same
uploads.idle_timeout_secs. - A lifecycle rule is still recommended:
{"Rules": [{"ID": "abort-incomplete-multipart", "Status": "Enabled", "Filter": {"Prefix": ""},
"AbortIncompleteMultipartUpload": {"DaysAfterInitiation": 2}}]}
Its delay, counted from initiation, exceeds the longest upload. MinIO
expires incomplete uploads on its own (stale_uploads_expiry).
Conditional writes
replaced in the audit and the reservation marker rely on
If-None-Match: * and If-Match of PutObject. Behind an endpoint_url (and
on AWS with cross_instance_reservation), the server measures them at
startup, on the key <prefix>/.craftfilegate-upload-precondition.<16 hex>,
deleted afterwards; the verdict holds until the next restart. MinIO: honoured.
Common keys on S3
| Key | On S3 |
|---|---|
home_dir (mount) | a prefix: the key is <prefix>/<home_dir>/<path>, empty parts omitted |
create_home (mount) | no effect: a prefix exists as soon as a key is under it |
refuse_upload_over_directory | true: an upload to a name that is also a directory is refused; cost: a one-key listing when each upload opens |
cross_instance_reservation | a marker object <lock_prefix><name>, put with If-None-Match: * |
lock_prefix | the name of these markers |
stale_partials | the sweep of abandoned multipart uploads |
case_insensitive | false by default: keys are case-sensitive |
hidden_stores | not applicable: an object only appears once complete |
Performance and limits
- Each upload in progress holds an 8 MiB buffer (Memory); its writes are sequential.
- An upload outside the root pays one listing per parent level.
- The client addresses the bucket in the URL path (path-style), AWS included.
- 10,000 parts at most per object, about 78 GiB at 8 MiB per part.
- On a versioned bucket, the self-test leaves two empty versions and one delete marker per startup.
What you will see
| When | Line |
|---|---|
| startup | INFO S3 backend initialized, fields bucket, prefix, endpoint |
| startup, self-test | INFO S3 upload precondition self-test, fields backend, verdict = honoured, probe_key |
| sweep | INFO aborted an abandoned multipart upload: ..., fields bucket, key, upload_id, age_secs |
| shutdown | INFO uncompleted S3 multipart uploads aborted before the shutdown completes, field aborted |
WebHDFS (Knox)
The webhdfs backend reads HDFS, read-only, through the WebHDFS of
Apache Knox: HTTPS, a service account over Basic, on behalf of the
user (doAs). Knox carries Kerberos to the cluster. It also serves an
HttpFS or a WebHDFS without Kerberos; Hadoop 2.8 and later.
[[backends]]
name = "datalake"
type = "webhdfs"
url = "https://knox.example.com:8443/gateway/default"
auth = { type = "basic", username = "svc-sftp", password_file = "/run/secrets/knox-password" }
ca_bundle = "/etc/craft-file-gate/knox-ca.pem" # the private CA of Knox
[[roles]]
name = "analystes"
[[roles.mounts]]
backend = "datalake"
mount_path = "/datalake"
home_dir = "/data/projets"
acl = [{ path = "/", rights = ["read", "list"], recursive = true }]
alice reads /datalake/2026/x.csv: the request targets
<url>/webhdfs/v1/data/projets/2026/x.csv?op=OPEN&doAs=alice.
The keys
All keys: Reference [[backends]] type = “webhdfs”.
The connection to Knox opens in 5 s at most; HTTPS_PROXY is not followed.
Common keys: Backends.
The mount on WebHDFS
| Key | On WebHDFS |
|---|---|
home_dir (mount) | an HDFS directory; . and .. are never resolved |
create_home (mount) | no effect |
acl (mount) | read and list: the rest has no effect |
With doAs, the user name is made of A-Z a-z 0-9 . _ -, 64
characters at most.
hidden_stores, refuse_upload_over_directory, stale_partials,
cross_instance_reservation and lock_prefix do not apply: nothing is
written. case_insensitive: leave it out, HDFS compares byte by byte.
Prerequisites
The right to name a user (doAs), with impersonate = true.
Who grants it depends on what is behind url:
Behind url | The setting |
|---|---|
| Knox (the intended case) | on the Knox side: in the topology, the identity assertion with impersonation enabled and hadoop.proxyuser.svc-sftp.users, .groups, .hosts; the topology exposes WEBHDFS and authenticates svc-sftp over Basic (ShiroProvider on LDAP/AD). On the cluster side, hadoop.proxyuser.knox.*, usually set by the Knox installation |
| Direct WebHDFS | svc-sftp declared as a proxy user in core-site.xml (below) |
| HttpFS | the same rights, under httpfs.proxyuser.svc-sftp. in httpfs-site.xml |
<property>
<name>hadoop.proxyuser.svc-sftp.hosts</name>
<value>sftp-gateway.example.com</value> <!-- where it connects from -->
</property>
<property>
<name>hadoop.proxyuser.svc-sftp.groups</name>
<value>sftp-users</value> <!-- on whose behalf it acts -->
</property>
At startup, the server runs a GETFILESTATUS of the HDFS root as the
service account, without doAs, 5 s at most.
What the backend does
| Operation | WebHDFS request |
|---|---|
stat, parent existence | GETFILESTATUS, one request per path |
| listing | LISTSTATUS_BATCH, page by page (startAfter); LISTSTATUS in one response on a cluster older than 2.8 |
| download, resume, read at an offset | OPEN?offset=&length=, by ranges |
upload, mkdir, rename, deletion | none: read-only |
- Read-ahead: 10 MiB read in 32 KiB SFTP packets cost 3
OPENwith the default; a read that skips only asks for its size. - Listings: read to the end, capped at 10,000 pages, 1,000,000 entries and 64 MiB per response.
- Redirects: a
307from Knox is followed once, only to the same origin (scheme, host, port) asurl. - A
SYMLINKentry is shown as a link, and cannot be read; HDFS compares names byte by byte, and so does the ACL.
Performance and limits
- Knox is the bottleneck: each
stat, page and range is an HTTPS request it relays to the cluster. The sessions of a backend share one HTTP client. timeout_secscovers a whole range: on a slow link, raise it rather than loweringread_ahead_bytes.- Small files each cost one
GETFILESTATUSand oneOPEN: this backend is for browsing HDFS, not for everyday storage. - The body of a Knox error is never copied into a log: only the exception name and the code.
The configuration file
One file, named at launch
craft-file-gate --config /etc/craft-file-gate/config.toml
-c is the short form. Only [auth] is required, with at least one
role and one door: [sftp], or [api] with [admin]. All the keys: the
reference.
| Section | What it sets | Page |
|---|---|---|
[server] | what all doors share: shutdown, sessions per user, atomic writes, probe port | Shutdown, Uploads, Admin API and console |
[sftp] | the SSH door, its algorithms, its bans | SFTP door, Bans |
[auth] | who gets in, and under which name | Authentication |
[[users]], [[roles]] | local accounts, roles and their mounts | Users, roles and mounts |
[[backends]] | the storages | Backends |
[uploads], [tcp_keepalive] | the timeouts | Timeouts, Uploads |
[log], [telemetry] | logs, audit trail, traces | Logs, Telemetry |
[admin], [api] | the admin console, the admin API, the file API, their bans | Admin API and console, Bans |
[reload] | hot reload | Hot reload |
[security] | how trusted files are judged | Trusted files and TLS |
Splitting: users_file and roles_file
Users, roles and backends can live in separate files. These files are
reloaded at runtime; config.toml almost not.
[auth]
users_file = "users.toml" # the [[users]]
roles_file = "roles.toml" # the [[roles]] and the [[backends]]
| Where to write | Effect |
|---|---|
[[users]], [[roles]], [[backends]] at the root of config.toml | read at startup |
the same under [auth] ([[auth.users]], …) | same; the root wins if both exist |
auth.users_file, auth.roles_file | read again on every change to the file |
| none of these | users.toml and roles.toml next to config.toml are used if they exist |
users_file and inline [[users]] do not add up: when users_file is
given, only that file is read.
Relative paths
A relative path to a file the server reads (host_keys, users_file,
roles_file, public_key_file, the persist_file keys, the pair and the
directory of [admin.tls]) starts from the directory of config.toml,
not from the launch directory: a configuration and its files travel
together, including in a container. The other paths (root of a local
backend, log.dir, the files of a WebHDFS backend) start from the launch
directory: write them as absolute paths.
The environment
A CRAFT_FILE_GATE_* variable replaces the key it overrides, after the
file is read, at startup and on every reload: the list is in
Environment variables. For a secret,
prefer the _FILE form: a file read once, judged as a
trusted file.
The binary’s tools
hash-password and verify-password read neither the configuration nor
the secrets; healthcheck reads the listeners of --config, nothing else.
| Command | Effect |
|---|---|
craft-file-gate hash-password | reads a password on stdin, writes its argon2id hash; --user <name> writes a [[users]] block; --profile picks the cost (see Authentication) |
craft-file-gate verify-password '<hash>' | reads a password on stdin; exits 0 if it produces this hash, 1 otherwise, 2 if the hash is unreadable |
craft-file-gate config check <file>, config schema | check a configuration without starting, write its JSON schema: Checking a configuration |
craft-file-gate healthcheck | queries /livez where the configuration serves it: probes_listen (HTTP), else control_listen, else admin.listen (HTTPS under [admin.tls]), read from --config (default /config/config.toml) and the environment, a 0.0.0.0 address queried on loopback; without that file, http://localhost:8080/livez; --url for another address; exits 0 on 200; for an image’s HEALTHCHECK |
What you will see
| When | Line |
|---|---|
| a declared file is read | INFO auth source loaded, fields kind (roles, users) and path |
| an undeclared file is found next to it | INFO found next to the config and used; declare it explicitly to pin it |
| a variable replaces a key | INFO config override from env, fields env and value; (value hidden) for a secret |
Checking a configuration
config check applies the startup validation to a file, without starting:
no port opened, no file created, no secret read, no network reached.
craft-file-gate config check /etc/craft-file-gate/config.toml
craft-file-gate config check config.toml --format json
craft-file-gate config schema > config.schema.json
craft-file-gate config explain sftp.ban.max_failures
craft-file-gate config effective /etc/craft-file-gate/config.toml
config check
| Option | Effect |
|---|---|
<file> | the configuration; its users_file, roles_file and trust files are read and judged as at startup |
--format text | default: one line per finding, <level> [<key>] <message> (see <link>) |
--format json | {"valid": ..., "findings": [{"level", "key", "message", "rule", "doc"}]} |
| Field | Meaning |
|---|---|
level | error: startup would refuse; warn: a startup WARN; info: a startup line, or a step left to startup, not checked offline: ... |
key | the key, written as in the reference (roles[].mounts[].backend) |
rule | the spec rule cited by the message |
doc | the key’s section in this guide |
| Exit | Meaning |
|---|---|
0 | no error, possibly some warn |
1 | an error, the first one, as at startup |
2 | unreadable file, or malformed command |
Environment variables count as at startup. A secret (host key, _FILE,
password_file, variable) is judged present, with its owner and mode; its
content is not read. Left to startup, reported as info: the JWKS, the
Kubernetes API, a storage probe, the upstream SFTP, the log files, the OTLP
exporter.
A CI job or an AI assistant writes the configuration, runs config check --format json, fixes each error with the help of key and doc, and
starts again until exit 0.
config schema
The JSON schema (draft 2020-12) of config.toml, users_file and
roles_file: types, defaults, bounds, allowed values, description of each
key, x-reload (hot reload), x-env (variables), x-doc (link). An
unknown key is refused; none is required.
A first line #:schema ./config.schema.json hands it to Taplo (Even
Better TOML): completion and underlining in the editor, as in
config.example.toml. The schema ships with each release and lives in the
images at /usr/share/craft-file-gate/config.schema.json;
config schema > config.schema.json writes it next to the configuration.
config explain <key>
One key: type, default, allowed values, hot reload (⟳), variables, effect,
spec, link to its section. --format json for a machine; an unknown key
exits with 2 and the closest ones.
config effective <file>
The configuration applied: the file, its linked files, the environment
variables, the defaults of the reference;
--format toml (default) or json. Every secret and every password hash
shows as ***, as does the identifier that goes with a secret
(access_key_id) and any value of a key named like a secret or under
headers. Offline like config check, which runs first: a refused
configuration exits with 1 and its finding.
Authentication
Two sources, three doors
| Door | Credential | Source |
|---|---|---|
| SFTP | password | local account, or JWT token under the sentinel name |
| SFTP | public key | local account |
| File API | Authorization: Basic | local account, password |
| File API | Authorization: Bearer | JWT token |
| Admin console and admin API | Authorization: Bearer | session token (local account, password), JWT token, or static token (see Admin API and console) |
A connection is authenticated by a single source, which gives its roles.
A local account alice and the sub alice of a token are two distinct
users.
Choosing the methods
[auth.methods]
jwt = { enabled = false }
local = { enabled = true, password = true, pubkey = true }
A keys-only deployment turns the password off:
local = { enabled = true, password = false, pubkey = true }
The keys: [auth.methods].
Local accounts
[[users]]
username = "alice"
password_hash = "$argon2id$v=19$m=19456,t=2,p=1$..." # craft-file-gate hash-password
authorized_keys = ["ssh-ed25519 AAAA... alice@poste"]
authorities = ["utilisateurs"] # their roles
The keys: [[users]]. In users_file,
they change without a restart.
Hashes
echo -n "mot-de-passe" | craft-file-gate hash-password # the hash only
echo -n "mot-de-passe" | craft-file-gate hash-password --user alice # a [[users]] block
The password is read on stdin, never as an argument. The hash goes to stdout, its parameters to stderr. The Hash tools tab of the admin console does the same in the browser, without sending the password.
Profile (--profile) | argon2id parameters | Memory per check |
|---|---|---|
owasp-min (default) | m=19456,t=2,p=1 | 19 MiB |
rfc9106-low-mem | m=65536,t=3,p=4 | 64 MiB |
A check costs what the stored hash says: the server never re-hashes. Keep a single profile per file: an unknown name then costs exactly the time of a wrong password, and cannot be guessed from the response time.
The argon2 formats ($argon2id$, $argon2i$, $argon2d$) are accepted.
To migrate from another server without knowing the passwords,
allow_sha512_crypt = true or allow_bcrypt = true under [auth.methods] local
also accepts these formats: Reference [auth.methods].
Each time the file is loaded, the line password hashes in use counts the
accounts per format, and so does the metric
craftfilegate_local_users{hash_format}: when a format drops to zero,
remove its flag.
local.max_argon2_memory_kib (19456: the m of the owasp-min profile)
names at load time every account whose hash asks for more; that account
still connects: re-hashing it with the default profile brings it under the
budget.
The check runs on a bounded thread pool, outside the runtime: see Password hashing.
The public key
The client offers a key, then signs if it is in its authorized_keys.
Offered keys do not count: an SSH agent that tries five before the right
one is not banned; a connection that ends without authenticating, after a
declined key, counts one failure.
The signature algorithms of a key
[sftp.algorithms] user_key sets the algorithms allowed to everyone (see
The SFTP door); ssh-rsa (SHA-1) is not in it by
default. A role can allow others to its users:
[[roles]]
name = "partenaire-ancien"
user_key_algorithms = ["ssh-rsa"] # in addition to user_key
[[roles.mounts]]
backend = "disque"
home_dir = "/partenaire"
The exception applies to every user who holds this role, never to a single user. Only the roles of the server’s files grant it, not the roles reached through the authorization service.
JWT
The jwt method is enabled by default.
[auth.jwt]
secret = "remplacez-moi" # or CRAFT_FILE_GATE_JWT_SECRET_FILE
algorithm = "HS256"
username_path = "/sub" # the user's name
authorities_path = "/groups" # their roles
| Door | Where to put the token |
|---|---|
| SFTP | login name jwt (auth.jwt_sentinel_username), the token as password |
| File API, admin API | Authorization: Bearer <token> |
One key source, your choice: secret (HS256, HS384, HS512),
public_key_file (the issuer’s PEM public key: RS256, RS384,
RS512, ES256, ES384) or jwks_url (the key set of an identity
provider, each key stating its algorithm). The keys:
[auth.jwt].
A token is accepted if its signature verifies, if exp is present and in
the future, and if nbf, when present, is past (60 s of leeway for both).
With issuer or audience, the iss or aud claim must be present and
equal. The user name is read at username_path, the roles at
authorities_path (an array of strings).
Plugging in an identity provider (jwks_url)
[auth.jwt]
jwks_url = "https://idp.example.com/.well-known/jwks.json"
algorithm = "RS256"
jwks_refresh_interval_secs = 3600
issuer = "https://idp.example.com/"
| Moment | What the server does |
|---|---|
| startup | validates the file, opens the log, then loads the key set, before opening the doors |
| each request | verifies the token against the key its kid header names |
every jwks_refresh_interval_secs | reloads the set; a key the provider added is learned at that moment |
| refresh with no answer | keeps the cached keys, which go on verifying tokens |
The cache age is the gauge craftfilegate_jwks_cache_age_seconds: alert
on it (see Metrics).
What you will see
| When | Line |
|---|---|
| each load of the accounts | INFO password hashes in use, one field per format |
| JWKS set loaded | INFO JWKS cache refreshed, field keys; then JWKS background refresh started |
| connection accepted | audit connection_accepted, auth_method = password, pubkey, jwt, basic or bearer; per key, signature_algorithm |
SFTP door and SSH algorithms
An example
[sftp]
listen = "0.0.0.0:2222"
host_keys = ["/etc/craft-file-gate/host_ed25519"]
[server]
max_sessions_per_user = 10
[sftp] is optional: without it, the SFTP door is off.
Create the host key once, before the first start:
ssh-keygen -t ed25519 -f /etc/craft-file-gate/host_ed25519 -N ""
The [sftp] keys
All keys: Reference [sftp]; timeouts: Timeouts.
[sftp.ban], [sftp.rate_limit]: Bans. In
[server], shared by all doors:
max_sessions_per_user (concurrent sessions of the same name, across all instances with [cluster]),
shutdown_grace_period_secs (Shutdown),
[server.hidden_stores] (Uploads).
The door advertises the password and publickey methods. It serves only
the sftp subsystem, once per connection: no shell, no exec, no port
forwarding.
Host keys
| Type | Format |
|---|---|
| ed25519, ECDSA (P-256, P-384, P-521), RSA | OpenSSH (ssh-keygen) or PEM (ssh-keygen -m PEM, PKCS#8) |
Each loaded key is reported with its SHA256:... fingerprint, the one that
ssh-keygen -l gives and that a client sees on its first connection. GET /admin/config and the Configuration tab of the console give the same ones.
generate_host_key creates the key only if its file is missing, in
0600, and gives its fingerprint. An existing file is read as is: the key
does not change from one start to the next. Keep it for a single instance on
persistent storage; in Kubernetes, mount a Secret (see
Kubernetes secrets).
A table entry chooses the algorithms a key advertises, with the
[sftp.algorithms] syntax applied to the host_key list:
[sftp]
listen = "0.0.0.0:2222"
host_keys = [
"/etc/craft-file-gate/host_ed25519",
{ path = "/etc/craft-file-gate/host_rsa", algorithms = ["rsa-sha2-512", "rsa-sha2-256"] },
]
An algorithm is advertised only if a loaded key can sign with it.
Algorithms
[sftp.algorithms] sets the algorithms offered, per category. Each list
starts from a modern default and is edited as in OpenSSH:
| Syntax | Effect |
|---|---|
["+name"] | appends name to the end of the default |
["-name"] | removes name from the default |
["a", "b"] | replaces the default with this list |
The defaults, list by list: Reference [sftp.algorithms].
kex,ciphers,macsandhost_keyare negotiated before the user is known: they apply to the whole server.user_keygives the algorithms a user key may sign its connection with (OpenSSH’sPubkeyAcceptedAlgorithms). A role can allow others for its own users only (user_key_algorithms, see Authentication).- The
ext-info-*andkex-strict-*markers are always added tokex. - The
server-sig-algsextension advertises to the client thehost_keylist in force, then theuser_keyalgorithms that no host key signs.
Weak algorithms
diffie-hellman-group1-sha1, diffie-hellman-group14-sha1,
diffie-hellman-group-exchange-sha1, aes*-cbc, hmac-sha1,
hmac-sha1-etm@openssh.com and ssh-rsa (RSA signed with SHA-1) are in
no default. Added with +, they are offered. diffie-hellman-group-exchange-sha256 can also be added, outside the
default.
An old client
# An old JSch client: neither curve25519 nor ML-KEM
[sftp.algorithms]
kex = ["+diffie-hellman-group14-sha1"]
| The client does not know | Add |
|---|---|
| curve25519, ML-KEM | kex = ["+diffie-hellman-group14-sha1"] |
ssh-ed25519 as a host key | an ECDSA P-256 host key in host_keys (ssh-keygen -t ecdsa -b 256 -f host_ecdsa -N ""), rather than an RSA one as ssh-rsa |
rsa-sha2-* for its user key | user_key_algorithms = ["ssh-rsa"] on its role only |
The sessions list of the console shows, for each session, the negotiated
algorithms and those that are weak (weak_algorithms).
What you will see
| When | Line |
|---|---|
| startup, each host key | INFO loaded host key, fields path, key_type, fingerprint |
| startup | INFO SSH algorithms offered, fields kex, ciphers, macs, host_key, user_key |
| connection accepted | INFO new SSH connection |
| session opened | INFO session established and registered, fields session_id, auth_method, signature_algorithm, total_sessions |
| ordinary end of a connection | INFO SSH session ended: <reason>, fields peer, username, client_version |
Uploads and atomic writes
Atomic writes (hidden stores)
By default, an upload writes into the destination file, like
HiddenStores off in ProFTPD: a reader can see a partial file while the
transfer lasts.
With enabled = true, the upload writes into an in-flight file in the
same directory, then renames it to the destination at the end. The
destination appears, or changes, only once the file is complete.
[server.hidden_stores]
enabled = true
prefix = ".in."
extension = "."
All the keys: Reference [server.hidden_stores].
The in-flight file is named <prefix><name>.<token><extension>:
.in.rapport.csv.3f9a1c04b7e25d68. with the defaults. The token, 16
hexadecimal characters, is unique to each upload. A .in.* pattern matches
them all.
enabled | During the transfer | Failed upload |
|---|---|---|
false | the file grows under its real name | the partial file stays under its real name |
true | the in-flight file grows next to it; the destination is untouched | the in-flight file is deleted; the destination is untouched |
In-flight files are listed and read like the others. They concern the
local and sftp (proxy) backends; an S3 object appears only once its
upload is finished, with no file next to it.
Per backend, per mount
The same table can be set at three levels. Each key resolves on its own:
the mount wins over the backend, which wins over [server.hidden_stores],
which wins over the defaults. SFTP and REST resolve the same value.
[server.hidden_stores] # the whole server
enabled = true
[[backends]]
name = "disque"
type = "local"
root = "/srv/sftp"
hidden_stores = { enabled = false } # this backend writes in place
[[roles]]
name = "depot"
[[roles.mounts]]
backend = "disque"
mount_path = "/depot"
home_dir = "/depot"
hidden_stores = { enabled = true } # except this mount
acl = [{ path = "/", rights = ["write", "list"], recursive = true }]
Leftovers of an interrupted transfer
A process killed in the middle of an upload leaves its in-flight file
behind. The server deletes it later, when an upload goes through the same
directory, once the file is older than grace_secs.
[uploads.stale_partials]
grace_secs = 900
age_check = true
All the keys: Reference [uploads.stale_partials].
The same table can be set per backend (stale_partials = { grace_secs = 1800 }
in its [[backends]]), key by key on top of [uploads.stale_partials].
On an S3 backend, it sets the cleanup of abandoned multipart uploads
(see S3).
A leftover is deleted only if its name has the exact form the server
writes, if nobody holds its lock, and if its age exceeds grace_secs on
the storage clock. At the slightest doubt, it stays. Each deletion leaves
an INFO line.
The instances that share a storage have the same
uploads.idle_timeout_secs: the threshold is computed from that of the
instance that sweeps.
Two uploads to the same destination
The first upload to open keeps the destination until it ends. A second one, from the same account or another, through the same door or another, is refused as soon as it opens, before sending a byte.
A stuck upload (client suspended by Ctrl+Z) is taken over by a retry
from the same account on the same instance, after
uploads.takeover_idle_secs without data.
Between several instances, a lock set on the storage, next to the
destination, carries the same rule. It is refreshed every 20 s; the lock of
a process that is gone is taken over after 2 x idle_timeout_secs + 60 s
(120 s by default). cross_instance_reservation and lock_prefix:
Reference [[backends]].
What the lock is on each type: Local,
S3, SFTP proxy.
Names that start with .craftfilegate-upload or with the lock_prefix
belong to the server: you can list and read them, not write to them.
With a single instance, cross_instance_reservation = false is enough.
Capping the size of a file
max_file_mb, on a mount, caps the size of each uploaded file
(1 MB = 1,048,576 bytes). It is not a space counter.
[[roles.mounts]]
backend = "disque"
home_dir = "/{username}"
max_file_mb = 500
| Value | Effect |
|---|---|
| absent | no cap |
N | each uploaded file is at most N MB, judged before the storage (in REST, on Content-Length before the body); what an upload over the cap had written is removed |
0 | no upload under this mount, read-only in effect |
What you will see
| When | Line |
|---|---|
| startup and reload, per backend | INFO stale in-flight files: ..., fields backend, grace_secs, grace_from, age_check, idle_timeout_secs |
| a leftover deleted | INFO removed a stale in-flight file: ..., fields backend, path, age_secs, clock, lock |
| lock of a gone process taken over | INFO took over a stale upload reservation: ... (local), took over the marker of an upload ... (S3), took over the lock file of an upload ... (proxy) |
| storage clock skew | metric craftfilegate_stale_partials_clock_skew_seconds{backend} |
Bans and rate limits
An example
[sftp.ban]
max_failures = 5
window_secs = 300
ban_duration_secs = 600
whitelist_ips = ["10.0.0.0/8"]
[admin]
listen = "0.0.0.0:8080"
[admin.ban]
max_failures = 10
trusted_proxies = ["10.0.0.5"] # the reverse proxy in front of the API
[api]
enabled = true
[api.ban]
max_failures = 5
trusted_proxies = ["10.0.0.5"] # the same one: a single listener
Without the section, the door bans nothing.
One list per door
| List | Section | Doors |
|---|---|---|
sftp | [sftp.ban] | SFTP |
api | [api.ban] | File API, download tickets, file explorer |
admin | [admin.ban] | Admin API and admin console |
The three lists are independent: an address banned on the file API still opens the admin console, and the other way around.
The keys of a ban
The same under [sftp.ban], [api.ban] and [admin.ban], read at startup.
All the keys: Reference [sftp.ban],
[api.ban],
[admin.ban].
What counts as a failure
A failure is a credential presented and judged wrong, whatever the account tried.
| Door | Counts |
|---|---|
| SFTP | wrong password or unknown name; JWT token judged wrong; key signature refused; a connection that ends without authenticating after declined keys (one failure per connection, not per key) |
| File API | wrong or malformed Basic; empty Bearer or token judged wrong |
| Admin console and admin API | wrong static token; empty Bearer, JWT or session token judged wrong; POST /admin/login: wrong password, unknown name or account without an admin role |
These do not count: a disabled method, an expired or not yet valid token, a refused download ticket, whatever it is, a refusal after an accepted credential (no role, session limit), a refusal by the rate limiter, the refusal of an address already banned. A successful connection does not reset the counter.
With [api] cors_origins = ["*"], any web page can make a visitor’s
browser send wrong credentials, which count toward the ban of the
visitor’s address: list the origins of your applications.
The life of a ban
| Moment | Effect |
|---|---|
max_failures failures within window_secs | the address is banned for ban_duration_secs; an ip_banned audit line |
with [cluster] | the failures the other instances hold in their window add up toward the threshold, read every second (≈ 1 s of delay); an unreachable instance stops counting after 5 s (2 × peer_timeout_ms + 1 s if longer) |
| on SFTP, at the verdict | the open SFTP sessions of the address are cut, transfers included |
| one more failure during the ban | the end moves back to ban_duration_secs after it |
| the end | the ban ends, with no line |
DELETE /admin/bans/{protocol}/{ip} | the ban is lifted |
An IPv4-mapped address (::ffff:10.1.2.3) is the same as 10.1.2.3, for
the ban, the whitelist and the proxies.
Lifting a ban
GET /admin/bans (Bans tab of the admin console) lists them; lifting one
requires the unban permission, {protocol} is sftp, api or admin:
curl -X DELETE -H "Authorization: Bearer $TOKEN" \
http://localhost:8080/admin/bans/sftp/203.0.113.7
A lift is a dated decision: it cancels the verdicts decided before it, never a later verdict. It does not reset the counter.
Sharing bans between instances
| Backend | Keys | Sharing |
|---|---|---|
| memory | none | each instance has its own bans, lost on restart; with [cluster], a lift is relayed to every instance, as with a file |
| file | persist_file | read at startup, published every 30 s, read again on a change and every reread_interval_secs; written atomically; on a shared volume (NFS, EFS) reading again is enough |
| ConfigMap | backend = "configmap", ban_configmap_name | watched through the Kubernetes API (feature k8s, a Role): Kubernetes |
The three lists can share a file or a ConfigMap. The merge keeps, per
address, the longest ban and the most recent verdict; a ban learned from a
peer cuts the SFTP sessions of the address, with no ip_banned line (the
one from the instance that decided it is enough). On shutdown, each list
is published, 5 s at most each.
The address of an HTTP client: trusted_proxies
The HTTP doors take the TCP peer as the address. When this peer is in the
trusted_proxies of their list ([api.ban] for the file API,
[admin.ban] for the admin API), they take the rightmost address of
X-Forwarded-For that is not a trusted proxy, failing that
X-Real-IP. The ban, the limiter, the share of the hashing queue and the
audit use this address. The SFTP door reads only the TCP peer.
Without [admin] control_listen, the file API and the admin API share a
listener: [api.ban] and [admin.ban] then carry the same
trusted_proxies.
Rate limits
A token bucket per address: burst connections or requests in a row,
then the rate per minute. Without the section, no limit.
[sftp.rate_limit]
connections_per_minute = 30
burst = 10
[api.rate_limit]
requests_per_minute = 600
All the keys: Reference [sftp.rate_limit],
[api.rate_limit].
The admin API has no rate limit. A rate refusal does not count toward the ban.
What you will see
| When | Line |
|---|---|
| a ban decided | audit ip_banned; INFO IP banned after auth failures |
| sessions cut by a ban | INFO cut the sessions of a banned address; audit session_end reason=banned |
| a peer publishes bans | INFO merged persisted bans from file |
Timeouts
What each timeout detects
| Timeout | Detects | Default |
|---|---|---|
sftp.login_grace_secs | an SSH connection that does not authenticate | 120 s |
uploads.idle_timeout_secs | a live client that stops sending during an upload (suspended, stuck) | 30 s |
uploads.min_rate_bytes_per_sec | a trickle upload | disabled |
[tcp_keepalive] | a peer gone with nothing in flight (machine powered off, NAT that forgot the connection) | ~60 s |
tcp_keepalive.user_timeout_secs | a peer gone while the server was sending it something; a client that stops reading | 65 s (SFTP), 35 s (HTTP) |
sftp.inactivity_timeout_secs | an SSH connection with no packet at all | 600 s |
sftp.keepalive_interval_secs | an SSH client frozen as a whole | disabled |
admin.header_read_timeout_secs | an HTTP request whose headers do not arrive | 10 s |
Uploads: [uploads]
[uploads]
idle_timeout_secs = 30
min_rate_bytes_per_sec = 1024 # 1 KiB/s on average; absent or 0: disabled
min_rate_grace_secs = 60
All the keys: Reference [uploads].
The idle timeout
The clock runs only while the server waits for the client. The time the server spends on its storage (slow S3, stalling NFS) never counts.
| Door | The clock |
|---|---|
| REST | restarts on each byte of the body |
| SFTP | one per write handle: starts at the OPEN, restarts on each complete WRITE of this handle, whatever the session does elsewhere; a packet must arrive whole within the timeout that follows its first byte |
When the timeout expires, the partial file is dropped (in-flight file deleted, S3 multipart aborted) and nothing is published.
A live upload reaches its file at least every 2 x
idle_timeout_secs: the cleanup of leftovers and the takeover of locks
rely on it, so give the same value to all instances that share a
storage. Raise it for clients that keep a file open without writing
(sshfs, a graphical client that asks for a confirmation, a very slow link:
at the default, a 32 KiB packet needs ~1.1 KB/s).
The minimum rate
Once min_rate_grace_secs of waiting have passed, an upload must have
received on average at least min_rate_bytes_per_sec, like Apache’s
RequestReadTimeout ... MinRate. The average runs from the start: a bursty
client keeps the credit of its bursts, a trickle is cut at the end of the
grace period. The bytes counted are those of accepted WRITE packets, not
the offset reached.
Below it, the upload is cut as for idleness. 1 KiB/s cuts a trickle and lets through a 10 KiB/s link that stalls.
TCP keepalive: [tcp_keepalive]
Set on each accepted connection, on the SFTP port and on the HTTP port.
All the keys: Reference [tcp_keepalive].
Network loss in the middle of a transfer: TCP_USER_TIMEOUT
Keepalive probes only a connection with nothing in flight. When the server
sends something to a client that is gone (the replies to the WRITE
packets of an SFTP upload, a download), TCP_USER_TIMEOUT cuts.
user_timeout_secs | SFTP port | HTTP port |
|---|---|---|
0 or absent | 2 x idle_timeout_secs + 5 s (65 s) | idle_timeout_secs + 5 s (35 s) |
N | N | N |
The connection is thus cut 5 s after the upload is abandoned. On Linux,
this timeout replaces count: an idle connection whose peer is gone is cut
at the first probe that exceeds it (around 70 s on SFTP). It also applies
to a client that stops reading: a REST download whose client is suspended
is cut at 35 s. A longer network loss (coverage loss, VPN reconnecting)
kills the transfer; raise user_timeout_secs if your clients go through
such losses.
Outside Linux, nothing is set.
SSH: [sftp]
login_grace_secs and inactivity_timeout_secs cut the connection (0:
never); keepalive_interval_secs (0: none) probes the client,
keepalive_max (3) cuts after that many probes without an answer. The
keys: [sftp].
The answer to an SSH keepalive counts as activity: with a keepalive
shorter than inactivity_timeout_secs, a live but idle session is no
longer cut for inactivity, only a client that no longer answers is.
That is why SSH keepalive is disabled by default.
HTTP: header read timeout
admin.header_read_timeout_secs (10 s, from 1 to 300,
CRAFT_FILE_GATE_ADMIN_HEADER_READ_TIMEOUT_SECS) bounds, on the port of
the admin API and the file API, the wait for the headers of a request:
from the connection, the TLS handshake or the previous response in
keep-alive. Beyond it, the connection is closed without a response. It
bounds neither a body nor a response.
The port serves only HTTP/1.1; over HTTPS, ALPN offers only http/1.1,
and HTTP/2 clients fall back to it.
What you will see
| When | Line |
|---|---|
| startup | INFO idle and liveness timeouts in force: ..., one value per field, and from_config: the keys set by the file |
| peer gone | INFO SSH session ended: the connection timed out: ... |
Logs
Two streams
| Stream | Target | Content |
|---|---|---|
| application log | everything but audit | startup, connections, warnings, errors |
| audit trail | audit | one line per operation, connection, refusal, ban, admin action |
Both go to stdout. With log.dir, they also go to two separate files.
[log]
[log]
level = "info"
format = "json"
audit = "all"
dir = "/var/log/craft-file-gate"
All the keys: Reference [log].
The volume of the trail: log.audit
log.audit removes only successes. A refusal, an error or an unknown
outcome is always written.
log.audit | Operation successes written | Ordinary connection_accepted, session_end |
|---|---|---|
all | all | yes |
changes | upload, delete, delete_recursive, rename, mkdir, rmdir, rmdir_recursive | no |
failures | none | no |
Always written, whatever log.audit: connection_rejected,
connection_rejected_summary, ip_banned, everything the admin API
writes, and the session_end of a cut session (administrator, shutdown,
ban, internal error).
The log level does not touch the trail: the audit target stays at info
whatever log.level.
Log files: log.dir
| File | Content | Retention |
|---|---|---|
craft-file-gate.log | the application log | retention_days (7 days) |
craft-file-gate-audit.log | the audit trail, and only it | audit_retention_days (90 days) |
- Rotation: on the first write of a new UTC day, and before a write
that would exceed
max_file_size_mb. The archive carries the moment it was opened,craft-file-gate.2026-09-27T00-00-00Z.log.gz, and holds only lines from that UTC day. - Retention: an archive goes once its day plus the retention has passed, judged on its name. The other files of the directory are not touched.
- Files in
0640,.gzarchives always complete; one directory per process.
In a container, stdout and the platform’s rotation (kubelet, journald,
Docker) remain the normal way; dir needs a volume there. The Helm chart
offers it: logFiles.enabled: true.
Repeated refusals: one summary line
An anonymous client that insists (scanner, misconfigured client) would
produce one line per attempt. Per door, address, reason and result,
the first refusal_summary_threshold refusals of a window of
refusal_summary_window_secs keep their line; the next ones are counted,
then summarized in one connection_rejected_summary line at the end of the
window, with suppressed, threshold and window_secs.
Only refusals that judged no credential are summarized:
missing credential, empty credential, malformed credential,
method disabled, verifier unavailable, banned, rate limit,
password checks saturated, password checks saturated for address,
shutting down. A wrong password always keeps its line.
Beyond refusal_summary_max_addresses tracked keys, the extra refusals
are summarized together (overflow=true).
SFTP session ends
The application line for the end of a session states its cause:
INFO line | Cause |
|---|---|
connection ended without a channel close — unregistering session | the client left |
SFTP channel closed — unregistering session | the client closed the channel |
session cut by an administrator — unregistering session | DELETE /admin/sessions/{id} |
session cut by the ban of its address — unregistering session | the address was just banned |
idle session cut by the server shutting down — unregistering session | shutdown, session with no transfer |
session cut by the server shutting down — unregistering session | shutdown, end of the grace period |
The trail carries one session_end line per session, with the same cause
in reason.
A connection that ends before the session (a client that does not speak
SSH, a negotiation with no common algorithm) writes INFO
SSH session ended: <reason>, with peer and client_version. These are
ordinary ends: alert on their number, not on each one.
What you will see
| When | Line |
|---|---|
| startup | INFO log filter in force, fields filter and source |
| startup | INFO audit trail filter in force (refusals and errors are always recorded), field audit |
startup with dir | INFO log files in force, both paths, the size and the retentions |
refusal_summary_* changed | INFO refusal summary settings reloaded |
| line that cannot be written (disk full) | metric craftfilegate_log_write_errors_total; one line on stderr at most per minute |
Log level and diagnostics
Changing the level at runtime
[log]
level = "debug"
Saving the file is enough: the level changes within moments, without a
restart, as for the users and roles files. On Unix, kill -HUP applies the
edit without waiting for the file watcher.
level | Effect |
|---|---|
trace, debug | more detail, for CraftFileGate only; dependencies (russh, hyper, the AWS SDK) stay at info |
info (default) | normal operation |
warn, error, off | fewer lines, for the whole process, dependencies included |
level takes a single value, not a directive string. The audit trail does
not depend on it: it stays at info.
In [log], only level, audit and the refusal_summary_* keys are
reloaded at runtime; the rest waits for a restart (see
Hot reload).
Who decides the level
| Source | Priority | At runtime |
|---|---|---|
RUST_LOG | the highest: replaces the whole filter | no: the level stays its own for the life of the process |
CRAFT_FILE_GATE_LOG_LEVEL | wins over the file | the file is no longer followed while it is set |
[log] level | the default | yes |
The first line of the log gives the filter in force and its source:
INFO log filter in force filter="info,craft_file_gate=info,audit=info,opentelemetry_sdk=off" source="config [log] level"
RUST_LOG, to go further
RUST_LOG gives the detail of a dependency, for example the russh trace.
It replaces the whole string, including the audit=info that the server
always adds: write it yourself.
RUST_LOG=warn,russh=trace,audit=info
| Directive | Target |
|---|---|
craft_file_gate=debug | the server (the crate name, with _) |
audit=info | the audit trail; audit=warn keeps only its refusals and errors |
russh=trace | the SSH protocol |
Understanding an authentication refusal
The audit line of a refused password does not say why: telling “unknown
user” from “wrong password” would reveal which accounts exist, and logs
travel far. The detail is at the debug level, for the operator only.
DEBUG local password authentication rejected username=alice reason="password does not match the stored hash"
DEBUG local password authentication rejected username=bob reason="no such user in the users file"
debug also gives the loaded accounts (local user store contents, field
usernames); info gives only their count (local user store loaded).
To check a hash without a server:
echo -n "mot-de-passe" | craft-file-gate verify-password '$argon2id$v=19$...'
Go back to info once the diagnosis is done. Other refusals carry their
cause in reason: Troubleshooting.
A runtime change writes INFO log level reloaded (level, source,
filter). Under level = "warn" or more severe, these announcements go to
stderr, prefixed craft-file-gate:: the announcement of a filter is never
silenced by it.
Telemetry
An example
[telemetry]
enabled = true
otlp_endpoint = "http://otel-collector:4318"
protocol = "http"
service_name = "craft-file-gate"
metrics = true
[telemetry]
All keys: Reference [telemetry].
The endpoint
protocol | otlp_endpoint | Spans sent to |
|---|---|---|
http | http://collector:4318 | http://collector:4318/v1/traces (and /v1/metrics) |
http | http://gw/otel/v1/traces | as is |
grpc | http://collector:4317 | as is |
https://is encrypted in both protocols; the collector’s certificate is checked against the system roots (see TLS trust roots).protocolis authoritative:OTEL_EXPORTER_OTLP_PROTOCOLandOTEL_EXPORTER_OTLP_TRACES_PROTOCOL, which some operators inject, do not change it.- The exporter reads
OTEL_EXPORTER_OTLP_HEADERS(authentication headers) andOTEL_EXPORTER_OTLP_TIMEOUT(10 s by default).
What is exported
The export receives the spans and events of CraftFileGate and of the audit
trail from info up, whatever the log level: a server set to warn
still exports its spans.
| Span | Attributes | Parent |
|---|---|---|
ssh_connection | peer, username | - |
auth_password, auth_publickey | username | ssh_connection |
sftp_session | session_id, username | the connection |
sftp_operation | operation, session_id, username, path, bytes | sftp_session |
api_request | method, path (as received, percent-encoded), username | - |
api_operation | operation, path (decoded), username | api_request |
authz_service_call | http.method, http.url, http.status_code | the caller |
The audit line of an operation is an event of its span; a failure sets the
span to status=error.
For a Tempo or Jaeger panel:
- a transfer that reaches
closehas twosftp_operationspans with the sameoperation(upload,download): the opening, then the commit, which alone carriesbytes. Count the spans that carrybytes; - an
rmdirof a non-empty directory has anrmdirspan then anrmdir_recursivespan, for a singlermdir_recursiveaudit line; - reads and writes have no per-packet span.
A path or a username that contains a control character or a Unicode
separator (U+2028, bidirectional controls) arrives escaped, in a visible
form (\u{2028}). The exact value is in the audit line.
Metrics over OTLP
With metrics = true, the server pushes its metrics every
metrics_interval_secs, in addition to GET /metrics, which is still
served; the name mapping: Metrics.
Sends, retries and shutdown
| Moment | Behavior |
|---|---|
transient refusal (http: 429, 502, 503, 504, no response; grpc: UNAVAILABLE, DEADLINE_EXCEEDED, …) | up to 4 attempts within the 10 s timeout; Retry-After honored |
| other refusal | a single attempt |
| collector unreachable | the batch is lost and counted, nothing stops; one line at the outage, one at the recovery |
| shutdown | last send of metrics and spans, 5 s at most |
Lost spans are counted on /metrics:
| Metric | Counts |
|---|---|
craftfilegate_otel_spans_ended_total | spans handed to the export |
craftfilegate_otel_spans_exported_total | spans accepted by the collector |
craftfilegate_otel_spans_export_failed_total | spans of a batch lost after its last attempt |
ended - exported - export_failed is what is still in flight (up to 2560
spans: the queue and the current batch), plus what a full queue rejected. A
gap that grows beyond that means lost spans.
increase(craftfilegate_otel_spans_ended_total[5m])
- increase(craftfilegate_otel_spans_exported_total[5m])
What you will see
| When | Line |
|---|---|
| startup, export off | INFO OpenTelemetry tracing disabled ([telemetry] absent or enabled = false); no spans are exported |
| startup, export on | INFO OpenTelemetry tracing enabled, fields endpoint (completed for http), protocol, service_name |
| startup, metrics | INFO OpenTelemetry metrics export enabled, fields endpoint, interval_secs |
| collector back after an outage | INFO OTLP collector recovered — telemetry export resumed |
| shutdown | INFO flushing OpenTelemetry spans, then OpenTelemetry shutdown complete |
Trust files and TLS
Trust files
These files decide who gets in and what they can do: anyone who can write them can grant themselves access. The server judges them at startup, and at reload for those that reload.
| File | Judged |
|---|---|
the file passed to --config | startup, reload |
sftp.host_keys | startup |
auth.users_file, auth.methods.local.users_file, or an adopted users.toml | startup, reload |
auth.roles_file, or an adopted roles.toml | startup, reload |
auth.jwt.public_key_file | startup |
admin.tls.cert_file, admin.tls.key_file | startup, reload |
admin.session.key_file | startup |
CRAFT_FILE_GATE_JWT_SECRET_FILE, CRAFT_FILE_GATE_ADMIN_BEARER_TOKEN_FILE | startup |
auth.password_file and ca_bundle of a webhdfs backend | backend construction |
SSL_CERT_FILE, SSL_CERT_DIR, if set | startup |
The file, and every directory crossed from / to reach it (symbolic links
included), is judged on its mode bits, its owner and its group. The server
then reads the file through the descriptor it judged.
Setting them up
| Setup | Verdict |
|---|---|
| owned by root, writable by its owner only, directories the same | accepted silently |
| owned by the server’s uid, writable by it only | accepted, reported at startup |
on a read-only mount (Kubernetes Secret or ConfigMap, :ro volume) | accepted; the directories are not judged |
/, /tmp (root, sticky bit) on the path | accepted |
The recommended setup:
chown root:root /etc/craft-file-gate /etc/craft-file-gate/*
chmod 755 /etc/craft-file-gate
chmod 644 /etc/craft-file-gate/config.toml /etc/craft-file-gate/roles.toml
chmod 640 /etc/craft-file-gate/users.toml /etc/craft-file-gate/host_ed25519
chgrp craft-file-gate /etc/craft-file-gate/users.toml /etc/craft-file-gate/host_ed25519
In a container, mount the configuration read-only (readOnly: true, :ro):
a Kubernetes Secret in 0440 with the pod’s fsGroup works as is.
[security]
All keys: Reference [security].
Read at startup. Keep it for a deployment where the server’s group contains only the server.
What the server creates itself
The process sets umask(077) at startup.
| Created by the server | Mode |
|---|---|
| generated host key, generated admin TLS key | 0600 (generated certificate: 0644) |
| directory of a generated key, bans file, temporary files | 0700 / 0600 |
| files uploaded by clients, home directories (local backend) | 0666 / 0777 masked by the umask inherited at launch |
| log and audit files | 0640 masked by the inherited umask |
Instances that share a bans volume run under the same uid.
Client paths
A client path is judged on its virtual form, then the backend resolves it under its root. On a local backend, a symbolic link never leaves the root: see Local.
TLS trust roots
This is outgoing TLS: S3 backend over HTTPS, jwks_url, authorization
service, OTLP collector over https://, WebHDFS backend. Incoming TLS for
the console is set in [admin.tls] (see
Admin API and console).
The server embeds no root: it reads the system store.
| Variable | Content |
|---|---|
SSL_CERT_FILE | a PEM file, any number of certificates |
SSL_CERT_DIR | directories, separated by : |
When set, either variable replaces the system store, it does not add to it. To trust an internal authority and the public roots, concatenate:
cat /etc/ssl/certs/ca-certificates.crt ca-interne.pem > bundle.pem
Without a variable, the first file found among the usual paths wins
(/etc/ssl/certs/ca-certificates.crt, /etc/pki/ca-trust/extracted/pem/tls-ca-bundle.pem,
/etc/pki/tls/certs/ca-bundle.crt, /etc/ssl/ca-bundle.pem, /etc/ssl/cert.pem),
then the directories /etc/ssl/certs and /etc/pki/tls/certs. The
:latest and :alpine images carry a store; the :scratch image has none.
Mounting a set of roots
The Mozilla bundle maintained by the curl project works:
curl -fsSLo bundle.pem https://curl.se/ca/cacert.pem
Mount it at the default path (/etc/ssl/certs/ca-certificates.crt,
read-only), or anywhere with SSL_CERT_FILE; in Kubernetes, a ConfigMap is
enough, it is not a secret.
A deployment without outgoing TLS connections (local backend or SFTP proxy, local accounts or static JWT key, no OTLP) needs no root.
Deployment recommendations
| Topic | Recommendation |
|---|---|
| trust files | read-only, owned by root; a Secret or a ConfigMap mounted readOnly: true |
| secrets | the _FILE forms rather than variables, readable in /proc/<pid>/environ |
| uid | a dedicated uid per instance, NoNewPrivileges=yes, a seccomp filter (SystemCallFilter=@system-service, seccompProfile: RuntimeDefault) |
| pod | the Helm chart applies the restricted profile: see Kubernetes |
| admin port | control_listen on an internal interface; in a container, behind a NetworkPolicy (the Helm chart sets one, and a separate ClusterIP Service) |
| OpenAPI description | [api] openapi, off by default: without authentication, it describes the whole API |
| console TLS | [admin.tls], certificate provided or generated |
| brute force | the three ban lists and the two limiters (see Bans) |
| audit trail | exported in full off the host; authentication refusals are at INFO, a WARN filter loses them |
| JWT | regular rotation of the signing keys; issuer and audience set |
Operating
| I want to | Page |
|---|---|
| install: image, Compose, binary | Docker |
| deploy in a cluster, with several pods | Kubernetes, Secrets |
| see sessions, cut one, lift a ban, read the logs | Admin console and API |
| give web access to the files | File explorer |
| know who did what | Audit: reading the trail |
| monitor and alert | Metrics |
| change the configuration without a restart | Hot reload |
| stop without cutting transfers | Shutdown |
| size memory and CPU | Performance and memory |
| understand an error | Troubleshooting |
Docker
ssh-keygen -t ed25519 -f ./host_ed25519 -N "" # the host key
sudo chown 65532 host_ed25519 # readable by the image's user
docker run -d --name craft-file-gate \
-v ./config.toml:/config/config.toml:ro \
-v ./host_ed25519:/config/host_ed25519:ro \
-v ./roles.toml:/config/roles.toml:ro -v ./users.toml:/config/users.toml:ro \
-v data:/data \
-p 2222:2222 -p 8080:8080 -p 8081:8081 \
craftogether/craft-file-gate:latest
config.toml follows the quick start
with what a container changes: listeners on 0.0.0.0, a key path
relative to the directory of config.toml, the console by account and password
(the doors are published), storage and accounts in their own files.
docker run --rm craftogether/craft-file-gate:latest config schema > config.schema.json
writes next to it the schema that #:schema names (Checking a configuration);
the image carries it at /usr/share/craft-file-gate/config.schema.json.
#:schema ./config.schema.json
[sftp]
listen = "0.0.0.0:2222"
host_keys = ["host_ed25519"] # /config/host_ed25519
[auth.methods]
local = { enabled = true } # the accounts in users.toml
jwt = { enabled = false } # no identity provider
[admin]
listen = "0.0.0.0:8080" # file API, file explorer
control_listen = "0.0.0.0:8081" # admin console, /admin, probes, /metrics
[[admin.roles]] # a [[users]] account with authorities = ["admins"]
name = "admins"
permissions = ["overview", "sessions", "kick", "bans", "unban", "config", "logs", "audit", "revoke", "grant"]
[admin.ban]
max_failures = 5
[api]
enabled = true
[api.ban]
max_failures = 5
[api.ui]
enabled = true
The first mount and its ACL: roles.toml ([[backends]], [[roles]]) and
users.toml ([[users]]), next to config.toml, are picked up without being declared.
#:schema ./config.schema.json
[[backends]] # roles.toml
name = "donnees"
type = "local"
root = "/data" # the data:/data volume
[[roles]]
name = "utilisateurs"
[[roles.mounts]]
backend = "donnees"
home_dir = "/{username}" # alice sees /data/alice as her root /
create_home = true
[[roles.mounts.acl]]
path = "/"
rights = ["read", "write", "list", "delete", "rename"]
recursive = true
#:schema ./config.schema.json
[[users]] # users.toml
username = "alice"
password_hash = "$argon2id$..." # output of hash-password (below)
authorities = ["utilisateurs", "admins"] # her files, and the admin console
The console: http://host:8081/, the file explorer: http://host:8080/files
(Console sessions, File explorer).
| Image | Contents | User | HEALTHCHECK |
|---|---|---|---|
:latest | static distroless, no shell | 65532 (nonroot) | yes |
:alpine | Alpine, with a shell for diagnostics | 65534 (nobody) | yes |
:scratch | the binary alone, no certificate store | none declared: set --user 65532 | no |
The image exposes 2222 (SFTP) and 8080; control_listen publishes a third one.
The HEALTHCHECK runs craft-file-gate healthcheck
on /config/config.toml. ENTRYPOINT is the binary alone, the command
--config /config/config.toml; the binary handles SIGTERM as PID 1.
- Mount read-only (
:ro) the files the server trusts: configuration, users, roles, host key, TLS pair (Security). - In the image,
/databelongs to its user, and a named volume inherits that; a host directory mounted in its place must belong to it:sudo chown 65532 data(65534for:alpine). :scratchhas no certificate roots: an outbound TLS connection (S3, JWKS, OTLP, Knox) needs a mounted bundle andSSL_CERT_FILE.- Subcommands pass through as is:
echo -n 'mot-de-passe' | docker run -i --rm craftogether/craft-file-gate hash-password. - Keep Docker’s stop timeout above the server’s:
docker stop -t 90. See Shutdown.
Docker Compose
services:
craft-file-gate:
image: craftogether/craft-file-gate:latest
restart: unless-stopped
stop_grace_period: 90s
ports:
- "2222:2222" # SFTP
- "8080:8080" # file API, file explorer
- "8081:8081" # admin console, /admin, probes, metrics
volumes:
- ./config.toml:/config/config.toml:ro
- ./roles.toml:/config/roles.toml:ro
- ./users.toml:/config/users.toml:ro
- ./host_ed25519:/config/host_ed25519:ro
- data:/data
volumes:
data:
The binary’s features
Each backend type and each door is a Cargo feature, all enabled
by default. The published binaries and images carry them all, plus k8s.
A binary built with fewer features refuses at startup the table
of a door or the type of a backend it does not carry, naming the feature.
| Feature | Backend or door |
|---|---|
backend-local | local; required (the server’s state, locks and bans live on the local disk) |
backend-sftp | sftp (proxy) |
backend-s3 | s3 |
backend-webhdfs | webhdfs |
door-sftp | the SFTP door, [sftp] |
door-rest | the REST file API and the file explorer, [api] |
k8s | bans and revocations shared through a ConfigMap, session key in a Secret; off by default in a custom build |
The console, the probes and /metrics have no feature: every binary
carries them.
Subcommands (hash-password, verify-password, healthcheck):
The configuration file.
Kubernetes
The Docker walkthrough under the chart: local storage on
a volume, the accounts in a Secret, the roles in config.toml.
kubectl create secret generic sftp-host-keys --from-file=host_ed25519
kubectl create secret generic sftp-users --from-file=users.toml
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: sftp-data }
spec: { accessModes: [ReadWriteOnce], resources: { requests: { storage: 10Gi } } }
EOF
helm install sftp deploy/helm/craft-file-gate -f sftp-values.yaml \
--set-file config.inline=config.toml
# sftp-values.yaml
hostKeys:
existingSecret: sftp-host-keys # /keys/host_ed25519
extraVolumes:
- name: data
persistentVolumeClaim: { claimName: sftp-data }
- name: users
secret: { secretName: sftp-users, defaultMode: 0440 }
extraVolumeMounts:
- { name: data, mountPath: /data } # the backend's root = "/data"
- { name: users, mountPath: /secrets/users, readOnly: true }
config.toml is the Docker one followed by the content of its roles.toml (the
chart mounts only config.toml), with host_keys = ["/keys/host_ed25519"]
under [sftp] and users_file = "/secrets/users/users.toml" under [auth].
Listeners stay on 0.0.0.0 (the probes reach the pod by its IP) and
[[admin.roles]] opens the console, since the chart sets the static token aside.
The volume belongs to the pod’s group (fsGroup 65532): create_home creates
alice/ there. Beyond one pod, the volume is ReadWriteMany.
The chart
| Value | Default | Effect |
|---|---|---|
replicaCount | 1 | number of pods; above 1, see Several pods |
config.inline | "" | the content of config.toml, rendered into the ConfigMap <release>-config; a change rolls the pods |
config.existingConfigMap | "" | or a ConfigMap of your own, key config.toml; exactly one of the two sources |
config.allowStaticToken | false | allow_static_token, written under [admin] in config.inline unless the configuration sets it; with existingConfigMap, write it yourself. See Authenticating |
extraVolumes, extraVolumeMounts, extraEnv, extraEnvFrom | [] | added as is to the pod and its container, after the chart’s: a volume for the root of a local backend, a Secret |
image.repository, image.tag, image.pullPolicy | craftogether/craft-file-gate, the chart’s version, IfNotPresent | the image |
hostKeys.existingSecret | "" | the Secret of the host keys, mounted in hostKeys.mountPath (/keys), files in 0440 with the pod’s fsGroup; see Secrets |
service.sftp.port, service.sftp.type | 2222, ClusterIP | the Service <release>, SFTP only |
service.admin.port, service.admin.type | 8080, ClusterIP | the Service <release>-api: file API, file explorer; absent with a config.inline without [admin] |
service.control.port | 8081 | control_listen: /admin, /metrics, probes; "" for a single door; ignored with a config.inline without [admin] |
service.control.type | ClusterIP | the type of the Service <release>-control: control port and exposed probe port |
service.probes.port | "" | probes_listen: /livez, /readyz, /health and /metrics on their own port, over HTTP; the probes go there |
service.probes.expose | false | this port on the Service <release>-control too: without it, probes and scraping go through the pod IP |
networkPolicy.enabled | true | an ingress NetworkPolicy on the pod |
networkPolicy.publicFrom | []: any source | sources allowed on SFTP and the admin port |
networkPolicy.controlFrom | - namespaceSelector: {}: any pod of the cluster | sources allowed on the control port and the probe port; at least one |
ban.backend | file | configmap to share bans between pods |
ban.configmapName | "": <release>-bans | the bans ConfigMap, created by the chart and given to the server through CRAFT_FILE_GATE_{SFTP,API,ADMIN}_BAN_CONFIGMAP_NAME |
stateDir.enabled, stateDir.path | true, /var/lib/craft-file-gate | an emptyDir for the bans file |
logFiles.enabled, logFiles.dir | false, /var/log/craft-file-gate | log files, in an emptyDir or logFiles.existingClaim |
passwordHashing.workers | "": requests.cpu | hashing threads; see Checking passwords |
passwordHashing.queue | "": 1024 | hashing queue |
resources | requests 100m, 64Mi; limits 500m, 256Mi | see Memory |
tls.enabled | false | probes over HTTPS, except on service.probes.port; [admin.tls] is set in config.toml |
terminationGracePeriodSeconds | "": the shutdown_grace_period_secs of config.inline plus 60 s, otherwise 90 s | the pod’s shutdown timeout, above the worst case of a shutdown; set it with an existingConfigMap that lengthens the timeout |
clusterSecret.*, adminRevocations.*, adminGrants.*, access.* | the session key and the revocations shared between pods: Secrets | |
cluster.enabled, cluster.port | "": as soon as replicaCount exceeds 1; 8083 | the channel between pods: Cluster |
cluster.minPeers, cluster.unreadyWhen | "": floor(replicaCount / 2); "": alert only | the detectors; unreadyWhen: a list, or auto (storage_alone, isolated_and_storage_down) |
- One Service per exposure: publishing SFTP through a
LoadBalancer(service.sftp.type) publishes neither the console, nor/metrics, nor the probes. - The kubelet probes come from the pod’s node, which most CNIs
let through despite the NetworkPolicy; otherwise, add the node CIDR
as an
ipBlocktonetworkPolicy.controlFrom. To restrict Prometheus and the ingress, list their namespaces innetworkPolicy.controlFrom.
Probes
/livez and /readyz (Probes and health) are
probed on the probe port when one is given, otherwise on the control
port, otherwise on the admin port.
A port for the probes
service.probes.port: 9090 sets [server] probes_listen: the probes and
/metrics get a port of their own, in plain HTTP, even under [admin.tls]
(One door or two). The kubelet and
Prometheus (per pod: PodMonitor or annotations) no longer need the
certificate or the console port; the port is on the Service
<release>-control only with service.probes.expose: true. An SFTP pod without
[admin] thus has its probes, with /readyz ready as soon as the SFTP door listens.
Pod hardening
The chart applies the restricted profile of the Pod Security Standards. Each
field can be overridden in podSecurityContext and securityContext; null
removes a field or a block.
| Setting | Value |
|---|---|
| user | runAsNonRoot: true, runAsUser, runAsGroup, fsGroup: 65532 |
| privileges | allowPrivilegeEscalation: false, privileged: false, capabilities.drop: [ALL] |
| seccomp | RuntimeDefault |
| root | readOnlyRootFilesystem: true: the server writes only to a volume |
| service account token | mounted with ban.backend: configmap, the shared Secret (clusterSecret) or the access ConfigMap (adminRevocations, adminGrants) |
| configuration, host keys | mounted readOnly: true |
The root of a local backend is a volume from extraVolumes. Host keys and the TLS pair come from a Secret
(cert-manager for TLS), never from an emptyDir: auto_generate of [admin.tls] does not apply under this chart.
Several pods
| What is shared | How | See |
|---|---|---|
| the pods find each other | as soon as replicaCount exceeds 1: headless Service <release>-cluster, port cluster.port reserved to the pods of the release (app.kubernetes.io/instance), certificate in the shared Secret | Cluster |
| console session key | the shared Secret (clusterSecret), entry admin-session.key | The pods’ shared Secret |
| console revocations, temporary access | the ConfigMap <release>-access (access.configmapName) | Sharing revocations |
| bans | ConfigMap (ban.backend: configmap), followed by a watch: a bit over 100 ms; or a file on a ReadWriteMany volume (persist_file), re-read every 5 s, eventual convergence | Bans |
| sessions, metrics, rate limits | nothing: each pod has its own; GET /admin/sessions lists those of the pod that answers, each with its pod_name (HOSTNAME); scope=cluster (the console) those of all |
Each Role covers a single name, without create: get, update on the Secret; get, update, patch
on each ConfigMap, watch on their collection for bans (helm template ... --show-only templates/rbac.yaml).
Sharing through a ConfigMap
Feature k8s (in the published images). The pod reads the ConfigMap before
serving; each publication is conditioned on the resourceVersion it read, and
a data.bans erased by kubectl edit is published again on the next one. Without
watch, propagation falls back to a re-read every 30 s; a
deployment that does not create the ConfigMap adds create on the collection.
[sftp.ban] # same under [api.ban] and [admin.ban]; the chart provides the name
backend = "configmap"
inotify on a shared node
fs.inotify.max_user_instances (128 by default) is counted per UID, across
the whole node. The server takes only one. When other pods have used them up and
you cannot tune the node:
[reload]
watch = "poll" # no inotify instance; a Secret update is seen within 5 s
See Hot reload.
Kubernetes secrets
| Secret | As a file | As an environment variable |
|---|---|---|
| whole configuration | config.toml mounted in /config | - |
| users, roles | users_file, roles_file | - |
| SSH host keys | host_keys | - |
| admin TLS pair | [admin.tls] cert_file, key_file | CRAFT_FILE_GATE_ADMIN_TLS_CERT, CRAFT_FILE_GATE_ADMIN_TLS_KEY |
| JWT secret | CRAFT_FILE_GATE_JWT_SECRET_FILE | CRAFT_FILE_GATE_JWT_SECRET |
admin token (fallback; forbidden by allow_static_token = false, which the chart sets) | CRAFT_FILE_GATE_ADMIN_BEARER_TOKEN_FILE | CRAFT_FILE_GATE_ADMIN_BEARER_TOKEN |
| S3 keys | credentials = { type = "static", ... } | AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY with type = "iam_role" |
| admin console session key | [admin.session] key_file | CRAFT_FILE_GATE_ADMIN_SESSION_SECRET_NAME: the shared Secret, filled by the pods (below) |
- A Secret mounted as a volume is updated by Kubernetes, which repoints the
..datalink: configuration, users, roles and TLS pair are reloaded at runtime. See Hot reload. - A Secret mounted with
subPathis never updated: avoidsubPath. - Mount Secrets and ConfigMaps
readOnly: true, as the chart does, never on a writableemptyDirorhostPath(Trust files and TLS). - A Secret passed through
envFrom(extraEnvFrom) is read only at startup: runkubectl rollout restartfor a new value. - The server does not read Secrets through the Kubernetes API: mount them. The only exception is the pods’ shared Secret (below).
A file rather than a variable
Mount the Secret as a file through the chart values, and name it with the
_FILE variable (the same for the admin token, with
config.allowStaticToken: true):
extraVolumes:
- name: jwt
secret: { secretName: sftp-jwt, defaultMode: 0440 }
extraVolumeMounts:
- { name: jwt, mountPath: /secrets/jwt, readOnly: true }
extraEnv:
- { name: CRAFT_FILE_GATE_JWT_SECRET_FILE, value: /secrets/jwt/jwt-secret }
The file is read once, at startup; /proc/<pid>/environ shows only a path.
See Environment variables.
users.toml is mounted the same way, named by users_file: see
Kubernetes.
The pods’ shared Secret
The pods of a release share the key that signs admin console sessions, so
that a pod accepts the token another pod issued. The chart creates an empty
Secret and gives the pods get and update on that name only. The first pod
that does not find the admin-session.key entry generates it and writes it;
the others read it. Each entry has its own life: the certificate of the
channel between pods lives there too (tls.crt, tls.key).
| Chart value | Effect |
|---|---|
clusterSecret.enabled | true by default, with [admin] or the channel between pods: the Secret, the Role (get, update on that name), the service account token; with [admin], CRAFT_FILE_GATE_ADMIN_SESSION_SECRET_NAME |
clusterSecret.name | <release>-cluster by default |
adminRevocations.backend | empty by default: configmap with the shared Secret, memory otherwise; configmap creates the access ConfigMap, where the revocations live, its Role (get, update, patch on that name), and sets CRAFT_FILE_GATE_ADMIN_SESSION_BACKEND, CRAFT_FILE_GATE_ADMIN_SESSION_REVOCATION_CONFIGMAP_NAME |
adminGrants.backend | empty by default: configmap with [admin] and the shared Secret or several replicas, memory otherwise; temporary accesses in the same ConfigMap, CRAFT_FILE_GATE_ADMIN_GRANTS_BACKEND, CRAFT_FILE_GATE_ADMIN_GRANTS_GRANT_CONFIGMAP_NAME; memory, or file without persist_file under [admin.grants], with several replicas makes rendering fail |
access.configmapName | the access ConfigMap, <release>-access by default |
With replicaCount above 1, clusterSecret.enabled: false requires a
key_file under [admin.session] in config.inline, and a shared key
requires shared revocations (adminRevocations.backend: configmap, or
persist_file under [admin.session]).
To change the key: remove the entry (kubectl patch secret ... --type=json -p '[{"op":"remove","path":"/data/admin-session.key"}]') then restart the
pods; open sessions are lost.
SSH host key
Create the key once, keep it in a Secret, mount it. Do not let the server
generate it (generate_host_key) in a pod: each pod, or each restart without
a volume, would have its own key, and clients would see “host key changed”.
ssh-keygen -t ed25519 -f host_ed25519 -N ""
ssh-keygen -l -f host_ed25519.pub # the fingerprint to give to clients
kubectl create secret generic sftp-host-keys --from-file=host_ed25519
hostKeys.existingSecret: sftp-host-keys mounts it in /keys (The chart),
and the configuration points to it: host_keys = ["/keys/host_ed25519"].
The fingerprint at startup (loaded host key, field fingerprint) and in
GET /admin/config is the one from ssh-keygen -l.
Several instances
The instances of a deployment find and talk to each other on a port of their
own, over TLS 1.3 in both directions, admitted by the certificate they
share. The Kubernetes chart sets all of this up as soon as replicaCount exceeds 1.
[cluster]
listen = "0.0.0.0:8083"
peers = "dns:craft-file-gate-cluster:8083"
cert_file = "/var/lib/craft-file-gate/cluster/tls.crt"
key_file = "/var/lib/craft-file-gate/cluster/tls.key"
All keys: Reference [cluster].
The shared certificate
Its SHA-256 fingerprint is the identity of the cluster: neither name nor date
is checked, expiry is reported by craftfilegate_cluster_cert_not_after_seconds.
To change it, remove the pair and restart every instance; during the
rollout, a peer with the old certificate is certificate_mismatch.
kubectl patch secret sftp-cluster --type=json \
-p '[{"op":"remove","path":"/data/tls.crt"},{"op":"remove","path":"/data/tls.key"}]'
kubectl rollout restart deployment/sftp
Detectors and withdrawal from service
After each round, each instance evaluates three detectors:
storage_alone (a backend down here that a reachable peer sees up: an
outage that everyone sees does not count), isolated (fewer than min_peers
reachable peers for isolated_after_secs) and isolated_and_storage_down
(isolated, and every non-local backend down here; a local backend, the
pod’s disk, does not count: without a non-local backend, it never
fires). By default they only alert:
WARN cluster detector active, craftfilegate_unready_detector{detector},
craftfilegate_cluster_isolated, checks.cluster of /admin/health.
Named in unready_when, they withdraw the instance from service: /readyz
not ready after unready_after_secs of an active detector, ready again after
ready_after_secs with none; /livez does not change. storage_alone only
withdraws a pod if another pod, in service and with no backend down,
sees each of its failed backends available, and only one pod at a time:
if every pod has an outage somewhere, none leaves.
[cluster]
min_peers = 1 # the chart: floor(replicaCount / 2)
unready_when = ["storage_alone", "isolated_and_storage_down"] # the chart: auto
isolated alone can withdraw every pod: cut of the cluster port,
loss of a majority of nodes, split into two equal halves, manual reduction
of the replica count. Only name it knowingly, and
never go below 2 × min_peers + 1 replicas without helm upgrade.
With the defaults, a storage outage is seen in 35 s at worst, the withdrawal follows 60 s later, and Kubernetes takes the pod out of the Service after 3 readiness probes at 10 s: about 2 min.
What you will see
At startup, INFO cluster channel listening (TLS 1.3, pinned on the shared certificate);
for each peer that answers, INFO cluster peer reachable; per status: craftfilegate_cluster_peers{status}.
The Instances, Sessions and Bans tabs of the console read each peer with your credential, which the peer judges itself (scope=cluster); a kick goes only to the peer that holds the session, a lift to every peer, except on a ConfigMap list; per call: craftfilegate_cluster_relay_total{route,result}.
A ban there is a guard against attempts, not a withdrawal of access: an address banned on a peer still reads it through another instance with a valid credential (a revoked session stays refused).
Admin API and console
[admin]
listen = "0.0.0.0:8080" # file API, file explorer
control_listen = "0.0.0.0:8081" # admin console, /admin, /metrics, probes
bearer_token = "changez-moi-en-production"
[api]
enabled = true
prefix = "/api/v1/files"
The keys: reference.
One door or two
| Configuration | Listener | What it serves | Runtime |
|---|---|---|---|
| without API | listen | console, /admin/*, /metrics, probes | admin (two admin-rt threads) |
API, without control_listen | listen | everything, and the API | main, shared with transfers |
control_listen | control_listen | /admin/*, /, /ui/*, /metrics, /health, /livez, /readyz | admin |
listen | file API, file explorer, /api/docs and /api/openapi.json | main | |
[server] probes_listen | probes_listen | /livez, /readyz, /health and /metrics, over plain HTTP, and nothing else; they leave listen and control_listen | admin |
With the file API, set control_listen: bans, kicks, /metrics and
probes stay reachable when transfers saturate the server. The Helm chart
does it (service.control.port). probes_listen gives probes and
Prometheus a port without TLS or console; without [admin], it is the only
HTTP listener (Kubernetes). The console
reads its health on /admin/health, served next to it in every case.
Authenticating
Each protected route requires Authorization: Bearer <token>.
| Token | Identity | Permissions |
|---|---|---|
a session token, returned by POST /admin/login (Console sessions) | the local account | those of the admin roles named by its authorities, reread at each request |
a JWT verified by [auth.jwt] | its name | those of the admin roles named by its roles claim (authorities_path) |
equal to bearer_token | <static-token> | all; a fallback, that allow_static_token = false disables (the Helm chart sets it) |
An admin role names permissions; an authority that carries its name grants them.
[[admin.roles]]
name = "support"
permissions = ["overview", "sessions", "logs"]
[[admin.roles]]
name = "admins"
permissions = ["overview", "sessions", "kick", "bans", "unban", "config", "logs", "audit", "revoke", "grant"]
The names of admin roles and those of [[roles]] are distinct. An account
that carries both kinds opens the console and the file doors; an account
that carries only admin roles opens only the console. The authorization
service never receives an admin role name, and no admin token opens a
file. Failures count in [admin.ban]
(Bans).
| Permission | Routes |
|---|---|
| none (any accepted token) | GET /admin/me, POST /admin/logout; POST /admin/session/renew (session token) |
overview | GET /admin/resources, GET /admin/status, GET /admin/health, GET /admin/doors |
sessions | GET /admin/sessions; GET /admin/events (session events) |
kick | DELETE /admin/sessions/{id} |
config | GET /admin/config, GET /admin/sessions/{id}/roles |
bans | GET /admin/bans; GET /admin/events (ban events) |
unban | DELETE /admin/bans/{protocol}/{ip} |
logs | GET /admin/logs/app, GET /admin/logs/app/stream |
audit | GET /admin/logs/audit, GET /admin/logs/audit/stream (each read is audited) |
revoke | GET /admin/revocations, POST /admin/revocations |
grant | GET, POST /admin/grants, DELETE /admin/grants/{id}, GET /admin/grants/candidates (accounts and mounts also require config, seen names sessions): a file role granted for a time ([admin.grants], reference); at until or on revocation, the SFTP sessions opened by the access and the REST transfers in progress under its mount are cut (on another instance: at its next reread); GET /admin/events (access events) |
| public | POST /admin/login, GET /admin/login, /, /ui/*, /health, /livez, /readyz, /metrics; /api/docs, /api/openapi.json with api.openapi |
Sessions
{"total": 1, "offset": 0, "limit": 50, "server_time": "2026-10-07T14:42:10.125Z", "inactivity_timeout_secs": 600,
"sessions": [{"session_id": "550e8400-e29b-41d4-a716-446655440000", "username": "alice", "remote_addr": "192.168.1.42:54321", "auth_method": "password",
"mounts": [{"mount_path": "/", "backend": "disque", "home_dir": "/alice", "acl": [{"path": "/", "rights": ["read", "write", "list", "delete", "rename"], "recursive": true}]}],
"connected_at": "2026-10-07T14:30:00Z", "idle_since": null, "last_sftp_op_at": "2026-10-07T14:41:25.402Z", "last_traffic_at": "2026-10-07T14:42:05.871Z",
"bytes_read": 1048576, "bytes_written": 524288, "sftp_ops": 412, "has_active_transfer": false, "pod_name": "craft-file-gate-0", "authorities": ["utilisateurs"],
"client_version": "SSH-2.0-OpenSSH_9.6", "algorithms": {"kex": "curve25519-sha256", "host_key": "ssh-ed25519", "cipher": "chacha20-poly1305@openssh.com",
"mac_client_to_server": "hmac-sha2-256", "mac_server_to_client": "hmac-sha2-256"}, "user_key": null, "weak_algorithms": [], "type": "sftp"}]}
| Field | Meaning |
|---|---|
mounts | one object per mount: mount_path, backend (its name), expanded home_dir, acl relative to the mount |
last_sftp_op_at | the last SFTP request, refusals included, keepalives excluded |
last_traffic_at | the last bytes from the client, keepalives included; the cut by inactivity_timeout_secs starts from there |
has_active_transfer | an open file; holds the shutdown during the grace period |
user_key | the SSH key: key_type, fingerprint, signature_algorithm |
weak_algorithms | the negotiated algorithms judged weak |
The list also carries REST transfers (upload, download) in progress
for long_request_threshold_secs (30 s by default, 0: all, reloaded at
runtime): "type": "rest" (an SFTP session has "sftp"), session_id,
username, remote_addr, method (PUT, GET), direction (upload,
download), path (the path seen by the user), mounts (mount_path,
backend), connected_at (its start), bytes, bytes_read,
bytes_written, bytes_per_sec (the average since the start).
| Request | Effect |
|---|---|
GET /admin/sessions?offset=&limit=&scope= | limit 50 by default, 1000 at most; scope=cluster: also the sessions of each instance, field instance, and peers (what each peer answered); local by default |
GET /admin/sessions/{id}/roles | the mounts as resolved at connection (effective.mounts, with roles, max_file_mb, create_home, hidden_stores) and the current definition of each role |
DELETE /admin/sessions/{id} | cuts the session ({"status": "disconnected"}) and removes it from the list and from the max_sessions_per_user count; a REST transfer is cut as if by a departed client, audit line reason="session ended: admin_kick"; held by another instance, that instance cuts it (instance) |
GET /admin/doors | the configured doors: name, readiness (accepting, not_accepting, hosted), started, sessions_open, activity |
Logs and streams
With [log] dir, GET /admin/logs/{app|audit} returns the last entries of
the current file, filtered by the server. See Logs.
Parameters: lines (500, 5000 at most; 4 MiB scanned at most), level
(ERROR to TRACE), since and until (RFC 3339), q (text, 256 bytes),
field=key:value (16 at most). 3 concurrent reads at most.
SSE streams (text/event-stream) follow activity live:
| Stream | Events |
|---|---|
GET /admin/events | session.opened, session.closed, session.kicked, session.revoked, ban.added, ban.lifted, lagged |
GET /admin/logs/{app|audit}/stream | entry, one per new retained entry; follows a rotation |
At most 8 open streams, 2 per identity; a stream ends at the exp of its
token or on its revocation, at the latest after 15 minutes.
Probes and health
| Route | Answers | Use |
|---|---|---|
/livez | 200 {"status": "alive"} as long as the process lives | liveness probe; craft-file-gate healthcheck |
/readyz | 200 if the SFTP port accepts and the main runtime takes a task within 500 ms, 503 otherwise and from the shutdown signal on | readiness probe |
/health | 200 always; status ok or degraded (authorization service unreachable), checks.auth_service, checks.data_plane (verdict of /readyz), nothing else | monitoring |
/admin/health | /health, with the version, active sessions and configured backends; permission overview | the console |
/admin/status | ready, degraded or not_ready, with each check (sftp, data_runtime, auth_service, jwks) and its detail | the console banner |
The console
http://serveur:8081/: a page that calls /admin/*. It loads nothing from
a third party. Its login: Console sessions.
| Tab | Content | Permission |
|---|---|---|
| Overview | status, version, sessions, authorization service | overview |
| Metrics | /admin/status banner, five-minute tiles, backends, top users | overview (top users: sessions) |
| Instances | with [cluster], each instance: status, version, sessions, bans, certificate expiry, health | overview |
| Sessions | one row per SFTP session and per long REST transfer, Type column (sftp, rest); mounts, activity, kick; with a peer, each instance (Instance column) | sessions, kick |
| Configuration | backends (secrets masked, storage clock), roles, host keys | config |
| Bans | current bans, lifting; with a peer, each instance (Instance column) | bans, unban |
| Logs | log and audit trail, filters, live follow; with [log] dir | logs, audit |
| Revocations | revoke the sessions of an account, those in force; when the door connects accounts | revoke |
| Temporary access | grant a file role for a time, live accesses and those ended in the last day, revocation; Temporary access; the Sessions view shows the accesses of a session | grant |
| Hash tools | hash-password and verify-password in WebAssembly in the page: the password does not leave the browser | none |
- Filters and sorts (
ip:,user:,role:,backend:,type:) live in the URL:/#sessions?q=user:bob&sort=-last_op. The token never appears there. - The page and its files (
/ui/*) carry a strict CSP (default-src 'none', everything from the same origin,'wasm-unsafe-eval'); a reverse proxy that sets its own allows at least as much.
What you will see
| When | Line |
|---|---|
| startup | INFO starting admin API server (HTTP) (or (HTTPS)) |
startup with [admin.tls] | INFO admin TLS certificate is valid, fields not_after, remaining_days; /metrics publishes the date |
Console session and revocations
A local account whose authorities name an admin role signs in to the admin
console with a username and password; the server gives it back a session
token that it signs itself.
[auth.methods]
local = { enabled = true, password = true }
[[users]]
username = "eric"
password_hash = "$argon2id$..."
authorities = ["admins"]
[[admin.roles]]
name = "admins"
permissions = ["overview", "sessions", "kick", "bans", "unban", "config", "logs", "audit", "revoke", "grant"]
[admin.ban]
max_failures = 5
[admin.session]
key_file = "/secrets/admin-session.key" # at least 32 bytes, the same on every instance
ttl_secs = 3600
max_age_secs = 43200
Sign-in counts its failures in [admin.ban], which it requires
(Bans), and goes through the
password hashing queue (Password hashing).
Signing in
| Request | Response |
|---|---|
POST /admin/login {"username": "...", "password": "..."} | token, to send as Bearer; username, expires_at, expires_in, auth_time, roles, permissions |
GET /admin/login | {"password": true} when the door signs in accounts |
POST /admin/session/renew (session token) | a new token, as long as max_age_secs since sign-in has not passed |
GET /admin/me (any token) | name, method, permissions, admin roles, expires_at, expires_in, auth_time |
The token carries only the account: each request re-reads the account in
effect ([[users]] or users_file) and recomputes its permissions. A reload
of users_file that removes a role, changes the password or removes the
account applies to the next request.
The session key
The key that signs tokens is the same on every instance: key_file, or under
Kubernetes secret_name
(Secrets). Without either,
the key is drawn at startup: sessions live with the process, on this instance
only. All the keys:
Reference [admin.session].
Revoking, signing out
| Request | Effect |
|---|---|
POST /admin/revocations {"username": "eric"} (permission revoke) | every session token of the account issued until then is refused, on all its tabs and devices; a new sign-in goes through. The response says lift: held (the shared state holds the revocation) or local_only (this instance only) |
POST /admin/logout | the same for the caller’s account; under a provider JWT or the static token, nothing is revoked (204) and the console forgets the token |
GET /admin/revocations | the revocations in effect; each one disappears max_age_secs after its issued_before |
Streams opened under a revoked token (/admin/events, log follows) are
closed.
Sharing revocations
[admin.session]
persist_file = "/var/lib/craft-file-gate/admin-revocations.json" # a file of their own
# or, on Kubernetes:
# backend = "configmap" # the chart sets it with the shared Secret
# revocation_configmap_name = "craft-file-gate-access"
reread_interval_secs = 5
All the keys: Reference [admin.session].
A shared key goes with shared revocations, and with synchronized clocks (NTP): a revocation is dated, and an instance whose clock runs ahead or behind applies it with that offset.
In the console
- Username and password when the door signs in accounts; “Sign in with a token” takes the static token or a provider JWT, and is the only form otherwise.
- The token lives in the page: reloading, opening a tab or closing the browser asks for sign-in again. The browser keeps only the theme and the table views.
- The session is renewed as long as the page is open, up to
max_age_secsafter sign-in; a closed tab lets it end atttl_secs. The header shows the account, its admin roles and the token’s end. - Logout calls
POST /admin/logout: one session signs out every tab and device of the account.
Temporary access
A temporary access gives a file role to a user name until a date, on top of the roles it already holds: a contractor for a delivery, support staff for a day. The name can be a local account or an identity brought by a JWT or the authorization service.
[[admin.roles]]
name = "support"
permissions = ["overview", "sessions", "grant"]
[admin.grants]
persist_file = "/var/lib/craft-file-gate/grants.json" # or backend = "configmap"
max_duration_secs = 604800 # 7 days
All keys: Reference [admin.grants].
- Who grants: an admin role that holds the
grantpermission, from the “Temporary access” tab of the console orPOST /admin/grants(Authenticating). Any role of[[roles]]can be granted, never an admin role. A reason is mandatory; the grant and its revocation are in the trail (grant_add,grant_revoke). - The name: the console offers the local accounts and the names that a file door authenticated in the last 7 days, on this instance. A name never seen is accepted with a warning: it opens nothing until someone connects under that name.
- The duration: 1 h, 8 h, 24 h, 48 h, 7 days or an end date, at most
max_duration_secs. - At the end (date reached or revocation): the SFTP sessions the access served and the REST transfers in progress under its mounts are cut, an upload in progress is abandoned; each following REST request is judged without it. A session opened before the grant does not see it.
- Several instances: the access lives in a file or the ConfigMap
<release>-access(one of the two with[cluster]), reread everyreread_interval_secs; another instance applies it and cuts it at its next reread. Clocks are assumed to be synchronized (NTP).
File explorer
[admin]
listen = "0.0.0.0:8080"
bearer_token = "changez-moi-en-production"
[api]
enabled = true
prefix = "/api/v1/files"
[api.ui]
enabled = true
path = "/files"
The page is at http://serveur:8080/files.
Options
Both keys are read at startup. All the keys:
Reference [api.ui].
What the page does
The page is a client of the REST file API, and of nothing else. It has
no rights of its own: what the ACL refuses, the API refuses, and the page shows
the server’s detail.
| Action | Request | Right |
|---|---|---|
| sign in | GET <prefix>?rights | - |
| open a folder | GET <folder>?list&offset=&limit=200, then ?rights | list |
| download | POST <file>?ticket, then the returned URL handed to the browser | read |
| upload | HEAD <file> (“Replace file?” if it exists), then a streamed PUT | write |
| new folder | PUT <folder>?mkdir | write |
| rename | POST <path>?rename=<destination> | rename |
| delete | DELETE <path> | delete |
- A button whose right is missing on the current folder is grayed out.
- Requests go one at a time, plus the upload in progress, to stay under
[auth] hash_per_addresswithBasic. - Uploads go one file after the other, with a progress panel (name, percentage, throughput, cancel).
- The list is paginated by 200, sortable by name, size or date, folders first.
- The current folder is in the URL (
#/rapports/2026): back, forward and links work.
Signing in
Username and password (Basic, local user), or “Use a token
instead” (Bearer, a JWT). The credential stays in memory only:
reloading the page signs you out. Only the theme is kept (craft-admin-theme).
Failures count for [api.ban].
Mounts
The page starts at /, the home that GET ?rights returns. A mount at /
shows its home_dir there; mounts under names show the
synthetic root that lists them.
| Entry | Display |
|---|---|
| mount point | storage icon, mount label; opens like a folder |
| synthetic directory | ordinary folder; upload, rename, delete grayed out |
Downloading: the ticket
A browser link carries no Authorization header: the page asks for
a ticket (this file, this user, ten minutes; see
The file API), then the browser
downloads as a stream. During those ten minutes, “Retry” and the browser’s
resume reuse the same URL.
The page’s files are public, under the same CSP as the console; the data goes through the API.
Limits
| Limit | Effect |
|---|---|
| no upload resume | an interrupted upload must be redone; a download restarted after ten minutes needs a new click |
| no preview or editing | no share link, no archive of a folder |
Basic behind a shared address (NAT) | users share hash_per_address; declare the proxy in api.ban.trusted_proxies, or sign in with a JWT |
Interface in English or French (?lang=fr, otherwise the browser’s
language); under 720 px, the actions move into a menu.
Audit: reading the trail
The audit trail says who did what, on which file, and whether it worked;
refusals and errors included. The vocabulary of the lines (actions, fields,
reason) is in the reference.
{"timestamp":"2026-10-07T14:41:25.402Z","level":"INFO","target":"audit","fields":{"message":"audit: operation succeeded","source":"sftp","action":"upload","session_id":"550e8400-e29b-41d4-a716-446655440000","remote_addr":"192.168.1.42","username":"alice","path":"/in/rapport.csv","backend":"disque","new_path":"","count":524288,"replaced":"no","removed":"","result":"success"}}
Where it goes
| Setting | Where the trail goes |
|---|---|
| default | stdout, mixed with the application log, one JSON line per event, target = audit |
[log] dir | in addition, craft-file-gate-audit.log (the trail alone) next to craft-file-gate.log |
[log] stdout = false | the files only |
[telemetry] | each operation line is also an event of the operation’s span |
Rotation, retention and level: Logs.
What it keeps
log.audit (all, changes, failures) only removes successes; repeated
anonymous refusals are summarized as connection_rejected_summary. See
The volume of the trail
and Repeated refusals.
Severity
| Level | Lines |
|---|---|
WARN | what only a credential holder produces: a denied, error or unknown operation; a refusal after an accepted credential; a refused admin action; ip_banned |
INFO | what any peer produces: the other connection_rejected; connection_accepted, session_end, the summaries, the successes |
An alert on the trail’s WARN lines therefore does not fire for an
anonymous scanner.
Filtering
In the console, Logs tab, source audit (permission audit): filters
by level, period, text and field:value; each read writes an
audit_read line. See Admin API and console.
# the admin console API
curl -H "Authorization: Bearer $TOKEN" \
'http://serveur:8081/admin/logs/audit?field=username:alice&level=WARN&since=2026-10-07T00:00:00Z'
# the files; without [log] dir: docker logs ... | jq -c 'select(.target == "audit")'
jq -c 'select(.fields.username == "alice")' /var/log/craft-file-gate/craft-file-gate-audit.log
| Question | Filter |
|---|---|
| everything an SFTP session did | .fields.session_id == "<id>" |
| overwritten files | .fields.action == "upload" and .fields.replaced == "yes" |
| ACL refusals | .fields.reason == "acl" |
| storage failures | .fields.result == "error" and .fields.backend == "<name>" |
| refusals and errors | .fields.result != "success" |
The text of a storage error is not in the trail: look for it in the
application log, at the same time, with the same path.
An incomplete trail
A line the server cannot write (closed pipe, full disk) is lost:
craftfilegate_log_write_errors_total counts it, stderr says so at most
once per minute, and the server keeps serving. Alert on
increase(craftfilegate_log_write_errors_total[5m]) > 0.
Metrics
curl http://serveur:8081/metrics # Prometheus text, no credential
# prometheus.yml
scrape_configs:
- job_name: craftfilegate
static_configs:
- targets: ['serveur:8081'] # probes_listen, else control_listen, else admin.listen
| Output | Format | Access |
|---|---|---|
GET /metrics | Prometheus text 0.0.4 | public; on [server] probes_listen or control_listen only, when they are set |
GET /admin/resources | the same snapshot in JSON, plus sessions, active bans, backends, thresholds | permission overview |
| OTLP export | the same numbers, one snapshot per interval | [telemetry] metrics = true, see Telemetry |
- Values are cumulative since startup, per pod. Compute a rate between two
reads (
rate()). - An unknown value is absent (
nullin JSON, no data point in OTLP), never0: the “Absent when” column says when. - The OTLP name is the Prometheus name without
_total.
Catalog
| Prometheus series | OTLP instrument | Type | Labels | Absent when | Meaning |
|---|---|---|---|---|---|
craftfilegate_connections_total | craftfilegate_connections | counter | SSH sessions opened after a successful authentication | ||
craftfilegate_connections_rejected_total | craftfilegate_connections_rejected | counter | refusals at the SFTP door, one per attempt (connection_rejected lines with source=sftp, summarized ones included) | ||
craftfilegate_sftp_operations_total | craftfilegate_sftp_operations | counter | op (21 values) | SFTP packets received, refusals included | |
craftfilegate_sftp_bytes_read_total | craftfilegate_sftp_bytes_read | counter | bytes served over SFTP | ||
craftfilegate_sftp_bytes_written_total | craftfilegate_sftp_bytes_written | counter | bytes accepted over SFTP | ||
craftfilegate_api_requests_total | craftfilegate_api_requests | counter | requests received by the file API, 429 and refusals included | ||
craftfilegate_banned_ips_total | craftfilegate_banned_ips | counter | ban verdicts of this pod, all doors | ||
craftfilegate_sftp_accept_errors_total | craftfilegate_sftp_accept_errors | counter | failed accept() on the SFTP port | ||
craftfilegate_log_write_errors_total | craftfilegate_log_write_errors | counter | log or audit lines not written | ||
craftfilegate_audit_refusals_suppressed_total | craftfilegate_audit_refusals_suppressed | counter | source (sftp, api, admin) | anonymous refusals summarized instead of written | |
craftfilegate_audit_refusal_summary_overflow_total | craftfilegate_audit_refusal_summary_overflow | counter | refusals that arrived with the summary table full | ||
craftfilegate_otel_spans_ended_total | craftfilegate_otel_spans_ended | counter | telemetry off | spans ended | |
craftfilegate_otel_spans_exported_total | craftfilegate_otel_spans_exported | counter | telemetry off | spans accepted by the collector | |
craftfilegate_otel_spans_export_failed_total | craftfilegate_otel_spans_export_failed | counter | telemetry off | spans lost | |
craftfilegate_admin_tls_cert_not_after_seconds | craftfilegate_admin_tls_cert_not_after_seconds | gauge | no TLS, unreadable date | expiry of the served certificate, Unix seconds | |
craftfilegate_admin_tls_cert_expiry_unreadable | craftfilegate_admin_tls_cert_expiry_unreadable | gauge | no TLS | 1 unreadable date, 0 read | |
craftfilegate_sftp_rate_limit_tracked_addresses | craftfilegate_sftp_rate_limit_tracked_addresses | gauge | no [sftp.rate_limit], or before its first sweep | addresses tracked by the limiter | |
craftfilegate_api_rate_limit_tracked_addresses | craftfilegate_api_rate_limit_tracked_addresses | gauge | no [api.rate_limit], or before its first sweep | same for the API | |
craftfilegate_local_users | craftfilegate_local_users | gauge | hash_format | no local store | local accounts per hash format |
craftfilegate_reload_watch_mode | craftfilegate_reload_watch_mode | gauge | mode (inotify, poll, sighup) | before arming | 1 for the active mode |
craftfilegate_jwks_cache_age_seconds | craftfilegate_jwks_cache_age_seconds | gauge | no JWKS, no successful load | age of the JWKS cache | |
craftfilegate_jwks_refresh_failures_total | craftfilegate_jwks_refresh_failures | counter | no JWKS | failed refreshes | |
craftfilegate_jwks_cached_keys | craftfilegate_jwks_cached_keys | gauge | no JWKS | keys in the cache | |
craftfilegate_grants_active | craftfilegate_grants_active | gauge | no [admin] | temporary accesses that grant their role now | |
craftfilegate_grant_sessions_cut_total | craftfilegate_grant_sessions_cut | counter | reason (expired, revoked) | no [admin] | SFTP sessions and REST transfers cut at the end of a temporary access they used |
craftfilegate_grants_store_local_only | craftfilegate_grants_store_local_only | gauge | no [admin] | 1: temporary accesses held in memory while a cluster peer answers | |
craftfilegate_cluster_peers | craftfilegate_cluster_peers | gauge | status | no [cluster] | peers of the channel between instances by status: reachable, unreachable, certificate_mismatch |
craftfilegate_cluster_relay_total | craftfilegate_cluster_relay | counter | route (health, sessions, bans, kick, unban), result (answered, refused, failed) | no [cluster] | operator requests relayed to a peer |
craftfilegate_cluster_cert_not_after_seconds | craftfilegate_cluster_cert_not_after_seconds | gauge | no [cluster], unreadable date | expiry of the shared certificate, Unix seconds; never checked | |
craftfilegate_cluster_reachable_peers | craftfilegate_cluster_reachable_peers | gauge | no [cluster] | peers reachable in the last round | |
craftfilegate_cluster_isolated | craftfilegate_cluster_isolated | gauge | no [cluster] | 1: fewer than min_peers peers reachable for isolated_after_secs (detectors) | |
craftfilegate_unready_detector | craftfilegate_unready_detector | gauge | detector (storage_alone, isolated, isolated_and_storage_down) | no [cluster] | 1 while the detector is active |
craftfilegate_withdrawn | craftfilegate_withdrawn | gauge | no [cluster] | 1: /readyz not ready because of a detector in unready_when | |
craftfilegate_stale_partials_clock_skew_seconds | craftfilegate_stale_partials_clock_skew_seconds | gauge | backend | nothing measured | storage clock minus server clock |
craftfilegate_password_hash_pool_places | craftfilegate_password_hash_pool_places | gauge | no local store | threads plus queue slots of the hashing pool | |
craftfilegate_password_hash_pool_in_use | craftfilegate_password_hash_pool_in_use | gauge | no local store | slots taken | |
craftfilegate_password_hash_pool_full_total | craftfilegate_password_hash_pool_full | counter | no local store | verifications refused because the pool was full | |
craftfilegate_jwt_refused_total | craftfilegate_jwt_refused | counter | no JWT | tokens refused by the verifier | |
craftfilegate_backend_operations_total | craftfilegate_backend_operations | counter | backend | backend unused | calls to the storage |
craftfilegate_backend_errors_total | craftfilegate_backend_errors | counter | backend | backend unused | storage failures (I/O, unreachable, credentials refused); not client refusals |
craftfilegate_backend_bytes_read_total | craftfilegate_backend_bytes_read | counter | backend | backend unused | bytes read from the storage |
craftfilegate_backend_bytes_written_total | craftfilegate_backend_bytes_written | counter | backend | backend unused | bytes written to the storage |
craftfilegate_backend_up | craftfilegate_backend_up | gauge | backend | before the first visit, probe off | 1 the storage answers the probe, 0 it no longer answers |
craftfilegate_backend_probe_seconds | craftfilegate_backend_probe_seconds | gauge | backend | same | duration of the last visit |
craftfilegate_local_symlink_refusals_total | craftfilegate_local_symlink_refusals | counter | backend | backend not local or unused | paths refused with symlink escape |
process_resident_memory_bytes | process_resident_memory_bytes | gauge | outside Linux, /proc silent | resident memory | |
process_peak_resident_memory_bytes | process_peak_resident_memory_bytes | gauge | same | peak resident memory | |
process_cpu_seconds_total | process_cpu_seconds | counter | same | CPU time, user + system | |
process_threads | process_threads | gauge | same | threads |
craftfilegate_connections_rejected_total counts per attempt, several per
connection: its ratio to craftfilegate_connections_total is not a failure
rate per connection.
Console thresholds
[admin.metrics_thresholds] only decides where the Metrics tab flags a tile;
the server does not alert. The keys and their defaults:
reference.
Alerts
groups:
- name: craftfilegate
rules:
- alert: CraftFileGateAdminCertExpiringSoon
expr: craftfilegate_admin_tls_cert_not_after_seconds - time() < 30 * 86400
for: 1h
- alert: CraftFileGateAdminCertExpiryUnknown
expr: craftfilegate_admin_tls_cert_expiry_unreadable == 1
- alert: CraftFileGateBackendErrors
expr: increase(craftfilegate_backend_errors_total[5m]) > 0
- alert: CraftFileGateBackendDown
expr: craftfilegate_backend_up == 0
- alert: CraftFileGateStorageAlone
expr: craftfilegate_unready_detector{detector="storage_alone"} == 1
- alert: CraftFileGateIsolated
expr: craftfilegate_cluster_isolated == 1
Alerting on a stale JWKS cache
- alert: CraftFileGateJwksStale
expr: craftfilegate_jwks_cache_age_seconds > 2 * 3600 # 2 x jwks_refresh_interval_secs
for: 5m
- alert: CraftFileGateJwksNeverLoaded
expr: absent(craftfilegate_jwks_cache_age_seconds) and on() craftfilegate_jwks_cached_keys >= 0
for: 5m
A stale cache keeps its keys: tokens signed by a key rotated since then are
refused. jwks_age_secs makes the same judgment for /admin/status.
Hot reload
vi /etc/craft-file-gate/users.toml # the server sees the edit and reloads
kill -HUP $(pidof craft-file-gate) # or: reload right away
[reload]
watch = "auto" # auto | inotify | poll
poll_interval_secs = 5 # period of the periodic reread
What is watched
config.toml always; users_file, roles_file and the [admin.tls] pair
when they are configured. The server watches the directory of each file: an
atomic replacement (sed -i, mv) and the switch of the ..data link of a
ConfigMap or Secret volume are seen. Events are grouped in 200 ms bursts, and
a single task re-reads. SIGHUP triggers a reload, even without watching.
What a reload does
It re-reads every watched file, whichever one changed.
| It applies | It keeps |
|---|---|
| the keys marked ⟳ in the reference | established SFTP sessions, with what they resolved at authentication |
users_file: users, hashes, keys, authorities | a refused file: nothing from it is applied |
roles_file: roles, mounts, ACL, backends, auth.authz_base_url client | the other keys: an edit is named, never applied |
| the admin TLS pair, if its bytes changed |
An authentication, or a REST request, reads the users in effect when it arrives. A re-read file is judged as at startup. A defined environment variable keeps the last word.
What a reload applies, key by key
The keys marked ⟳ in the reference; a key
without the mark is read at startup. The [[users]], [[roles]] and
[[backends]] entries are reloaded when they live in users_file or
roles_file; written in config.toml, they are read at startup. The
[admin.tls] paths need a restart; the content of the pair, however, is
re-read at every reload.
inotify or periodic re-read
All the keys: Reference [reload].
Both keys are read at startup.
- A single inotify instance per process, shared with the ban files;
pollfits where they are scarce (fs.inotify.max_user_instances, per UID and for the whole node). - The periodic re-read compares date, size, inode, ctime, owner and mode; an interval without change writes nothing.
- When the kernel event queue overflows (
fs.inotify.max_queued_events), every file is re-read. Repeated overflows mean that a neighbor is churning the configuration directory: move the configuration, or raise the limit.
What you will see
INFO line | Meaning |
|---|---|
hot reload armed (file watch + SIGHUP), ... (periodic re-read + SIGHUP) | the active mode; field watching |
file change detected, reloading configuration (or ... by the periodic re-read ..., SIGHUP received, ...) | a reload starts |
hot-reloaded users file, hot-reloaded roles file, log level reloaded, audit trail filter reloaded | what is applied |
inline config — roles not hot-reloadable | at startup: roles and backends in config.toml |
Shutdown
[server]
shutdown_grace_period_secs = 30 # wait for work in progress, all doors
docker stop -t 90 craft-file-gate # Docker
# Kubernetes: terminationGracePeriodSeconds: 90
The steps
SIGTERM or SIGINT (Ctrl-C) trigger the shutdown, even when the server is PID 1.
| Step | What happens | Bound |
|---|---|---|
| 1 | the SFTP door stops accepting; /readyz answers 503; no password check is admitted any more | - |
| 2 | the ban lists ([sftp.ban], [api.ban], [admin.ban]) are written to their persist_file | 5 s per list |
| 3 | server shutting down sent to every SFTP session; the admin/API door stops accepting and closes its streams | - |
| 4 | sessions without a transfer in progress are cut | - |
| 5 | wait for transfers in progress, if any; ends with the last one | shutdown_grace_period_secs |
| 6 | any session still there is cut | - |
| 7 | wait for the cut sessions to end (spans, last audit lines) | rest of the delay |
| 8 | wait for REST and admin requests in progress | rest of the delay |
| 9 | S3 multipart uploads still open are aborted | 5 s |
| 10 | quota cleanups and upload lock removals | 5 s |
| 11 | last OTLP export (metrics, spans) | 5 s |
| 12 | graceful shutdown complete, exit 0 | - |
Steps 5, 7 and 8 share a single delay, counted from the announcement. Without a transfer in progress, the shutdown does not wait for it.
Tuning the orchestrator
The worst case of a shutdown:
| Component | Default |
|---|---|
shutdown_grace_period_secs | 30 s |
| ban writes | 5 s per list ([sftp.ban], [api.ban], [admin.ban]) |
| backend work | 5 s |
| cleanups and locks | 5 s |
| OTLP export | 5 s |
| total | 50 s with one ban list, 60 s with all three |
The orchestrator’s shutdown delay must exceed it: docker stop -t 90,
terminationGracePeriodSeconds: 90, which the chart sets (grace period plus 60 s). A SIGKILL before the end cuts
transfers without an audit line, leaves S3 multipart uploads open and loses
unexported spans.
What you will see
Each step writes its INFO line: shutdown signal received,
ban files written out before the shutdown sequence, idle sessions disconnected (idle_kicked), waiting for active transfers to complete
(only if there are any), remaining sessions disconnected (force_kicked,
0 included), graceful shutdown complete. Cut sessions end
with session_end reason=shutdown_idle (step 4) or shutdown (step 6).
Performance and memory
| Question | Page |
|---|---|
| which memory limit to set | Memory |
| how many password verifications in parallel | Password hashing |
| what to watch in production | Metrics |
Memory
The formula
peak ≈ 12 MiB base (14.5 MiB with glibc)
+ 5.5 MiB if an S3 backend is used
+ hash_workers × m + 2 MiB password checks
+ sessions × 0.15 MiB
+ connections in the hash queue × 70 KiB at most hash_queue
+ local transfers × 4 MiB
+ S3 uploads × 26 MiB + S3 downloads × 8 MiB
+ addresses tracked by bans × 225 B
m is the largest argon2 m in the users file: 19 MiB with the
owasp-min profile (default of hash-password), 64 MiB with the
rfc9106-low-mem profile. The startup line gives this term:
peak_memory="1 thread × 19 MiB = 19 MiB".
See Password hashing.
The formula is an upper bound: it adds up peaks that do not all happen at the same moment. No measured case has exceeded it. File size does not enter it: transfers are streamed.
The terms are measurements on the published image (static musl binary,
x86_64), with an OpenSSH client. A connection in the hashing queue weighs
~42 KiB in Basic, a local download 2.3 MiB. At large scale with S3, uploads
dominate: 4 threads at m=65536 and 10 S3 uploads measure 415 MiB
(formula: 537). Nothing caps the number of concurrent S3 uploads;
max_sessions_per_user sets a bound per user.
A small deployment under 80 MB
30 users, ~900 files per day: a few concurrent clients.
[auth]
hash_workers = 1 # one check at a time: 1 × 19 MiB
hash_queue = 32 # a full queue: 32 × 70 KiB = 2.2 MiB
With hash-password hashes (profile owasp-min), and no
rfc9106-low-mem hash in the file:
| Variant | Formula | Fits under 80 MB |
|---|---|---|
| local backend | 45.5 MiB | yes |
| 2 threads, local backend | 64.5 MiB | yes |
| S3, one upload at a time | 64.7 MiB | yes |
| S3, 3 concurrent uploads | 116.5 MiB | no |
profile rfc9106-low-mem | 90.5 MiB | no: a single verification exceeds the budget |
| default queue (1024) | +70 MiB with a full queue | no |
glibc binaries
The Docker images (musl) give memory back after each peak. A glibc binary
(releases, glibc host) keeps the memory of the hashes in its arenas:
95 MiB retained after a burst (4 threads, m=19456), 19 MiB with
MALLOC_MMAP_THRESHOLD_=4194304. Set this variable (Environment= in the
systemd unit) and do not set MALLOC_ARENA_MAX, which raises the peak.
Kubernetes
| Setting | What it must cover |
|---|---|
limits.memory | the peak of the formula, for the expected number of concurrent transfers and sessions, plus a margin |
requests.memory | memory at rest and under ordinary load: base and sessions |
The chart ships requests.memory: 64Mi and limits.memory: 256Mi. With its
default (1 thread, derived from requests.cpu: 100m), 1000 SSH connections
queued with the rfc9106-low-mem profile reach 146 MiB. On a busy node, raise
the request toward the expected peak.
A pod without limits.cpu sees the node’s CPUs: see
Password hashing to set the pool size.
Watching memory
/metrics exposes process_resident_memory_bytes and
process_peak_resident_memory_bytes (Linux only), and the admin console plots
them live. See Metrics.
Not measured: REST transfers, SFTP proxy, active telemetry, thousands of users, aarch64 and Windows.
Checking passwords
[auth]
hash_workers = 2 # threads that compute the hashes
hash_queue = 1024 # attempts waiting for a thread
hash_per_address = 4 # slots one address can hold at once
Options
The three keys are read at startup; a reload does not change the
size of the pool. All the keys: Reference [auth].
How it works
An argon2id, bcrypt or sha512-crypt hash is pure computation: from 16 ms
(owasp-min) to more than a second (heavy argon2). Checks run
on dedicated threads (craft-file-gate-hash-0, -1… in ps -L), off the
runtime that serves transfers, the API and the probes.
| Situation | What happens |
|---|---|
| a thread is free | the attempt is checked at once |
| threads busy, room in the queue | the attempt waits its turn |
| threads busy, queue full | immediate refusal, with no account lookup, not counted for the ban |
the address already holds hash_per_address places | same refusal, whatever the load |
| client gone before its turn | the attempt is not hashed; its place is given back |
| address banned while waiting | the attempt is not hashed; banned-address response |
| server shutdown | no new attempt admitted; those on a thread finish |
An attempt takes its place before the account lookup: a known name and an unknown name get the same response. Rate limiters and bans (Bans) bound what a burst costs the queue.
Behind a reverse proxy, declare it in trusted_proxies: each client
counts at its own address. A NAT address goes in the door’s whitelist_ips
([sftp.ban], [api.ban], [admin.ban]), which is not capped. Console
sign-in (POST /admin/login) goes through the same pool.
Choosing the sizes
| Question | Answer |
|---|---|
| why at most 2 threads by default | each thread can hold a whole argon2 m; 2 threads at the owasp-min profile stay under 80 MB |
| how many connections per second | 2 threads: about thirty at the rfc9106-low-mem profile (68 ms), over a hundred at the owasp-min profile (16 ms) |
| why 1024 places | a place is a waiting connection (~70 KiB over SSH, ~42 KiB with Basic), not an argon2 allocation; 100 clients reconnecting together wait instead of being refused |
| wait of the last place | hash_queue / hash_workers checks: 6.5 s with 2 threads at the owasp-min profile; stay under sftp.login_grace_secs (120 s) |
| tight memory budget | a shorter queue, sized to the expected burst (hash_queue = 32); the pool holds hash_workers times the largest m in the file: Memory |
What you will see
At startup, one line gives each number and its source:
INFO password hashing pool started: passwords are checked on these threads, off the async runtime, ...
workers=2 workers_source="detected (8 CPUs ...), capped at 2 by default: ..."
queue=1024 queue_source="default (1024, ...)"
per_address=4 per_address_source="default (4 per address, ...)"
peak_memory="2 threads × 19 MiB = 38 MiB"
tokio_workers=8 tokio_workers_source="detected (std::thread::available_parallelism)"
tokio_workers: the threads of the async runtime, from the same detection with no
cap, or TOKIO_WORKER_THREADS.
A reload of users.toml that changes the largest m says so:
INFO password hashing peak memory changed with the users file.
Kubernetes
A pod without limits.cpu sees the node’s CPUs (32, 64). The cap at 2 protects
the default pool, not the tokio runtime. To follow what the pod requests,
pass requests.cpu through the Downward API (rounded up to the next integer):
env:
- name: CRAFT_FILE_GATE_HASH_WORKERS
valueFrom:
resourceFieldRef:
containerName: craft-file-gate
resource: requests.cpu
divisor: "1"
# same for TOKIO_WORKER_THREADS
An explicit value has no cap: requests.cpu: 8 gives 8 threads and
8 x m of memory. The Helm chart does it with passwordHashing.workers (empty:
requests.cpu) and passwordHashing.queue.
Troubleshooting
An index of errors: the exact word you see (startup message,
reason of an audit line, SFTP status, REST code) leads to the guide page
that defines the option at fault, then, in parentheses, to its table in the
reference, and to the identifier of the behaviour rule, to quote to support.
<...> marks a variable part. The full vocabulary of audit lines is in the
reference.
Startup and configuration
A broken setting refuses to start; a setting with no effect starts with a
WARN that names the key.
| Message | Where it is set | Rule |
|---|---|---|
--config is required when not using a subcommand, Configuration error: failed to read config file: <cause> | The configuration file | R-CONFIG-001 |
Configuration error: failed to parse config TOML: unknown field <key>, expected one of ... (at line <l>, column <c>) | the guide page of the named section (reference): the misspelled key | R-CONFIG-005 |
... `sftp.<key>` is now `server.<key>`: move it to the [server] table: shutdown_grace_period_secs, max_sessions_per_user or hidden_stores written under [sftp] | The keys of [sftp] ([server]) | R-CONFIG-015 |
[sftp]: this binary is built without the SFTP door (Cargo feature door-sftp) (or [api], REST, door-rest), no door is configured, nothing would be served: ..., this binary is built without any door ... | reference: [sftp], or [api] with [admin]; The binary’s features: a binary built without that door | R-CONFIG-016 |
server.probes_listen = <address> is also <key> — the probes need a port of their own | One door or two ([server]) | R-ADMIN-020 |
server.backend_probe.<key> = <n> is out of bounds: between <min> and <max>, server.backend_probe.timeout_secs = <n> is not below interval_secs = <m>: ... | Availability ([server.backend_probe]) | R-AVAIL-001 |
Configuration error: <section>: <cause>, <key> = <n> is out of range: accepted values are <min> to <max> | the guide page of the named section or key (reference) | R-CONFIG-010 |
Configuration error: failed to load local users: <cause> | Local accounts ([[users]]) | R-AUTH-004 |
duplicate username: <name> — each [[users]] entry needs its own username | Local accounts ([[users]]) | R-CONFIG-011 |
user <name> has a sha512-crypt password hash, which is not accepted — set [auth.methods] local.allow_sha512_crypt = true ... (or bcrypt); WARN local user's password_hash is the example published in the README or the example files ..., local user's argon2 hash asks for more memory per verification than auth.methods.local.max_argon2_memory_kib allows: ..., local users whose password hash has another format or other parameters than most of the users file: ..., local users whose password hash cannot be verified (...): ..., accounts still depend on a legacy password hash flag: ..., legacy password hash flag enabled but no account uses this format ..., password hash is not argon2id — ... | Hashes ([auth.methods]) | R-AUTH-004, R-AUTH-005, R-AUTH-007, R-AUTH-008, R-AUTH-010, R-CONFIG-012 |
Configuration error: failed to load roles/backends: <cause>; duplicate role name, role without [[roles.mounts]], mount to an unknown backend | The keys of a role and a mount, Keys of every backend ([[roles]], [[roles.mounts]], [[backends]]) | R-AUTH-028 |
unknown field backend (or home_dir, acl) under [[roles]]: a mount key written at the role level | The keys of a role and a mount ([[roles.mounts]]) | R-CONFIG-005 |
role has no ACL entry: access is deny-by-default, ... (WARN) | Writing an ACL ([[roles.mounts.acl]]) | R-ACL-003 |
role "<name>" holds ACL paths that name one path on this backend, whose ACL compares paths folded (case and Unicode normalization), with different rights: [...] | Name case, Case of ACL paths ([[roles.mounts.acl]], [[backends]]) | R-ACL-006 |
role missing required fields: role '<name>': mount '<p>': home_dir = "<h>" climbs out of the backend's root | Where the files go ([[roles.mounts]]) | R-AUTH-032 |
user '<name>': roles '<a>' and '<b>' both claim '<mount_path>' ... (at startup, or failed to reload roles file): conflicting mounts of a local user | Combining roles, Mount recipes ([[roles.mounts]]) | R-AUTH-030 |
role '<name>': user_key_algorithms: "<algorithm>" is not an algorithm this server supports; ... | Signature algorithms of a key ([[roles]]) | R-AUTH-019 |
no authentication methods enabled | Choosing the methods ([auth.methods]) | R-AUTH-001 |
no JWT key source, jwks_url with another source, unknown algorithm; first JWKS load failed, the first fetch from auth.jwt.jwks_url failed, ...; WARN auth.jwt.public_key_file is never read: ..., every JWT will be refused: auth.jwt.secret and auth.jwt.public_key_file are both set ..., [auth.jwt] is configured but auth.methods.jwt.enabled is false: ... | JWT, Connecting an identity provider ([auth.jwt]) | R-AUTH-020, R-AUTH-026, R-AUTH-025 |
auth: auth.hash_workers = 0 is refused: it must be between 1 and 1024 ... | Checking passwords ([auth]) | R-AUTH-014 |
unknown backend type "<type>": this binary knows <types> | Types, The binary’s features ([[backends]]) | R-CONFIG-013 |
sftp.host_keys: <path> does not exist. Create it (ssh-keygen ...) or set sftp.generate_host_key ..., sftp.host_keys: cannot read <path>: <error> (the key must be readable by the image’s user) | Host keys, Docker, Secrets ([sftp]) | R-SFTP-003 |
| a refused algorithm list (unknown name, mixed forms, empty list) | Algorithms ([sftp.algorithms]) | R-SFTP-006 |
[cluster] has no listen ..., ... has no peers ..., ... has no certificate source ..., ... sets both cert_file/key_file and secret_name ..., cluster.peers entry "<entry>" cannot be read: ..., cluster.secret_name is empty: ..., cluster.secret_name is set, but this binary was built without the k8s feature: ..., cannot listen on cluster.listen = <address>: <cause>, the cluster channel cannot start, refusing to boot: ..., cluster.<key> is set without cluster.<other> ..., cluster.peer_timeout_ms = <n> is out of bounds ..., cluster.listen = <address> is also <key> ..., cluster certificate (cert_file=<c>, key_file=<k>): <reason>; WARN [cluster] has nothing to add across instances: ... | Cluster ([cluster]) | R-CLUSTER-001, R-CLUSTER-002 |
cluster.<key> = <n> is out of bounds ..., cluster.unready_when names <detector>, but cluster.min_peers = 0 turns isolation off ..., ... but server.backend_probe.enabled = false ..., cluster.<key> = <n> is under <min> s (...): one storage verdict could withdraw or restore the instance, CRAFT_FILE_GATE_CLUSTER_MIN_PEERS = ... cannot be read ..., CRAFT_FILE_GATE_CLUSTER_UNREADY_WHEN = "..." cannot be read: ...; WARN a [cluster] detector key is set where it has no effect | Detectors and withdrawal from service ([cluster]) | R-CLUSTER-016 |
WARN cluster detector active (detector, reason), cluster detector inactive; withdrawn from service: ..., back in service: ...; /readyz 503 withdrawn_by; storage_alone active without withdrawal | Detectors and withdrawal from service: storage_alone, the storage of this pod alone (mount, node network); active without withdrawal: no pod in service and healthy (with no backend down) sees all its down backends available, or a pod with a smaller identifier goes first; isolated, the network between pods; isolated_and_storage_down, isolated and every non-local backend down: the node network; the reason is in /admin/health (checks.cluster) | R-CLUSTER-015, R-CLUSTER-016 |
<anchor>: [the directory ]<path> is <reason>: whoever can write there decides who gets in. ...; a file this server trusts can be rewritten: ... (WARN) | Trusted files: chmod go-w, or mount read-only | R-TRUST-008 |
<anchor>: <path> is writable by the server's own group (<gid>): ... | [security] ([security]) | R-TRUST-004 |
<anchor>: <path> cannot be resolved: <error> | Setting them | R-TRUST-007 |
private key readable by others (WARN) | Trusted files | R-TRUST-011 |
<X> and <X>_FILE are both set: one secret, two sources..., <X>_FILE is set but empty, ... is empty: an empty secret is no secret, ... is not UTF-8 text | A file rather than a variable (Environment variables) | R-CONFIG-009 |
<VARIABLE>="<value>" is refused: it must be a whole number...; env override ignored, the configured value stands, env override is empty (WARN) | The environment (Environment variables) | R-CONFIG-008 |
server.shutdown_grace_period_secs must be > 0 | Shutdown ([server]) | R-SHUTDOWN-004 |
admin requires a bearer_token or at least one role under [[admin.roles]] | Authenticating ([admin], [[admin.roles]]) | R-ADMIN-001 |
admin.listen and sftp.listen are the same address — they cannot share a port, admin.control_listen = <address> is also admin.listen — the control door needs a port of its own, the admin door cannot start, refusing to boot: ... cannot listen on <address> (port already taken) | One door or two ([admin]) | R-ADMIN-001, R-ADMIN-002 |
tls: cert_file and key_file must both be set, tls: enabled but no cert source ..., tls: cert_file and auto_generate are mutually exclusive; pair refused at startup: <cause> (cert_file=<path>, key_file=<path>) | Deployment recommendations ([admin.tls]) | R-ADMIN-001, R-ADMIN-003 |
[[admin.roles]] has an entry with a blank name, [[admin.roles]] names the role <name> twice, [[admin.roles]] <name> grants no permission: ..., unknown variant <permission>; [[admin.roles]] is set but no credential can open the admin door: ... | Authenticating ([[admin.roles]], [auth.methods], [auth.jwt]) | R-ADMIN-006 |
role <name> is also the name of an [[admin.roles]] entry: admin roles and file roles need different names (at startup, or failed to reload roles file, keeping old config); an admin role and a file role have names that differ only by case or spaces: ... (WARN) | Authenticating ([[admin.roles]], [[roles]]): rename one of the two | R-ADMIN-006 |
this user's keys open no door: its authorities name no role with mounts, only admin roles, ... (WARN) | Authenticating ([[admin.roles]]): remove authorized_keys, or give a file role | R-ADMIN-006 |
the admin door signs local accounts in with their passwords ... and has no [admin.ban]: ... | Console sessions ([admin.ban]) | R-ADMIN-022 |
admin.session.key_file <path> holds <n> bytes: a session key needs 32 at least ..., admin.session.key_file: ... writable by ..., admin.session.key_file <path>: <cause>; admin.session.ttl_secs = <n> is out of range ..., admin.session.max_age_secs = <n> is under admin.session.ttl_secs ..., admin.session.key_file and admin.session.secret_name are both set ..., admin.session.secret_name is empty ..., admin.session.secret_name is set, but this binary was built without the k8s feature ... | The session key ([admin.session]): head -c 32 /dev/urandom, chmod 400; The binary’s features | R-ADMIN-022 |
auth.jwt.issuer = "craft-file-gate" is the issuer of the console's own session tokens ... | [auth.jwt]: another issuer | R-ADMIN-005 |
admin session key generated for this process: ... (WARN), tokens refused by another pod or after a restart | The session key ([admin.session]): key_file or secret_name | R-ADMIN-022 |
admin.allow_static_token = false, and a static admin token is set by <source>: ...; the static admin token is enabled beside admin roles: ..., static admin token used: it is meant for break-glass only (WARN) | Authenticating ([admin]): remove the token from <source>, or allow_static_token = false | R-ADMIN-024 |
api.enabled requires [admin] section to be configured | REST file API ([api], [admin]) | R-REST-001 |
[api.ui] enabled = true requires [api] enabled = true: ..., [api.ui] path "<path>" is not usable, ... collides with ... | File explorer ([api.ui], [api]) | R-EXPLORER-001, R-EXPLORER-002 |
[api] openapi = true requires [api] enabled = true: ... | REST file API ([api]) | R-REST-012 |
admin.metrics_thresholds.jwks_age_secs = <age> is not above auth.jwt.jwks_refresh_interval_secs ...; WARN a metrics threshold is set where it has no effect: ... | Console thresholds ([admin.metrics_thresholds], [auth.jwt]) | R-METRICS-017 |
invalid telemetry otlp_endpoint, invalid telemetry protocol, invalid telemetry metrics_interval_secs; cannot build the OTLP <signal> exporter: ...; refusing to start | The endpoint, [telemetry] ([telemetry]): on_exporter_error | R-TELEMETRY-001, R-TELEMETRY-003 |
unknown ban backend ...: expected "file" or "configmap", sftp.ban: ..., api.ban: ..., admin.ban: ...; the file API has no ban list: set [api.ban], ..., the admin routes have no ban list: set [admin.ban], ..., [api.ban] is set where it has no effect: ... (WARN) | The keys of a ban, One list per door ([sftp.ban], [api.ban], [admin.ban]) | R-BAN-018 |
[api.ban] trusted_proxies [...] and [admin.ban] trusted_proxies [...] differ, while the file API and the admin routes share [admin] listen: ..., [admin.ban] trusted_proxies [...] is set and the file API has no [api.ban]: ..., [api.ban] trusted_proxies [...] is set and there is no [admin.ban], while ... share [admin] listen: ... | The address of an HTTP client ([api.ban], [admin.ban]): both lists with the same trusted_proxies, or [admin] control_listen | R-BAN-021 |
hidden_stores: prefix shaped like a lock, no affix, grace_secs out of bounds (... is too short: ...) | Atomic writes ([server.hidden_stores], [[backends]], [[roles.mounts]]) | R-HIDDEN-004, R-HIDDEN-005, R-HIDDEN-014 |
uploads.takeover_idle_secs is not below uploads.idle_timeout_secs: ... | Two uploads to the same destination ([uploads]) | R-RESERVE-017, R-CONFIG-010 |
WARN create_home is set on a mount whose backend is not local: ..., home_dir holds a marker that is not {username}: ..., follow_symlinks = true on this local backend: ... | Common keys on a local backend, The root and symbolic links, {username} | R-LOCAL-002, R-LOCAL-005, R-AUTH-032 |
lock_prefix invalid; cross_instance_reservation = true is not supported on a local backend on Windows | Reservation across instances ([[backends]]) | R-RESERVE-013, R-RESERVE-016 |
reload.poll_interval_secs = <n> is out of range: accepted values are 1 to 60; refusing to start with reload.watch = "inotify" and the inotify instance cannot be created (or a directory cannot be watched) | Hot reload ([reload]): watch = "auto" or "poll" | R-RELOAD-010 |
a [telemetry] metrics key is set where it has no effect, ... is set where it has no effect (WARN) | the guide page of the named key (reference): remove it | R-CONFIG-010 |
this secret is the example published in the README or the example files (WARN) | Secrets: change the secret | R-CONFIG-012 |
outbound TLS is configured but there are no CA certificates to verify it with, ... (WARN), cause: SSL_CERT_FILE points at <path>, which does not exist, so it is ignored | TLS trust roots (Environment variables) | R-TRUST-014 |
Hot reload
A file refused at reload keeps what is in force. See Hot reload.
| Message | Where it is set | Rule |
|---|---|---|
configuration edited but not applied until restart (keys=...) | What a reload applies, key by key (reference): restart | R-RELOAD-003 |
failed to reload users file, keeping old config, failed to reload roles file, keeping old config, failed to build backends registry from reloaded roles, keeping old config, failed to re-read the config for its log level, keeping the current one, [log] level is not a log level (accepted values: ...); keeping the log filter currently in force | Local accounts, The keys of a role and a mount, Keys of every backend, [log] ([[users]], [[roles]], [[backends]], [log]): the field e gives the cause | R-RELOAD-005 |
password hashing pool size edit not applied: the pool is sized once, at startup | Checking passwords ([auth]) | R-AUTH-014 |
[log] format edit ignored: ... | [log] ([log]): restart | R-AUDIT-029 |
[log] level edit ignored: <variable> is in force ... | Who decides the level ([log], Environment variables): the variable wins | R-CONFIG-007 |
admin TLS certificate reload failed, still serving the previous certificate; admin TLS certificate reload refused, still serving the previous certificate (ERROR) | What a reload applies, key by key, Trusted files ([admin.tls]): pair unreadable, mismatched or refused | R-RELOAD-008, R-TRUST-010 |
file hot reload cannot use inotify, falling back to re-reading the files every <n> s; the file watcher lost events: its event queue overflowed, ..., the file watcher failed and may have lost events ... | Hot reload ([reload]) | R-RELOAD-010, R-RELOAD-012 |
reload panicked, previous configuration kept; hot reload still armed (ERROR) | to report | R-RELOAD-013 |
Connection and authentication
reason of the connection_rejected lines, and what the client sees. The HTTP
doors answer an error in application/problem+json, field detail.
reason, code or detail | Where it is set | Rule |
|---|---|---|
absent, 401 invalid credentials (the file explorer: Invalid credentials; 429: Too many attempts); at DEBUG, local password authentication rejected (or public key), reason no such user in the users file, password does not match the stored hash, the stored hash could not be parsed, no local user store configured, user has no authorities | Local accounts ([[users]]): wrong password or unknown name | R-AUDIT-022, R-AUTH-006, R-EXPLORER-006 |
no authorized key offered | The public key ([[users]]): authorized_keys | R-AUTH-016 |
signature algorithm not allowed | Signature algorithms of a key ([[roles]], [sftp.algorithms]) | R-AUTH-018 |
missing credential, empty credential, malformed credential; 401 missing or invalid authorization header, invalid Basic auth, JWT not configured, local auth not configured | REST file API, Choosing the methods ([auth.methods]): send Basic or Bearer | R-REST-003 |
invalid token, expired token, wrong issuer, wrong audience, 401 invalid JWT | JWT ([auth.jwt]): issuer, audience, key | R-AUTH-022 |
verifier unavailable; ERROR JWKS background refresh failed ... (the cache keeps its keys) | JWT, Connecting an identity provider ([auth.jwt]) | R-AUTH-026, R-AUTH-027 |
no username claim, 401 token carries no username claim | JWT ([auth.jwt]): username_path | R-AUTH-024 |
method disabled, 401 authentication method disabled | Choosing the methods ([auth.methods]) | R-AUTH-003 |
no matching roles, 403 no roles resolved; role resolution, 503 roles could not be resolved, retry later; WARN authz service returned non-200, authz service call failed | Where roles come from, The authorization service contract ([[users]], [auth.jwt], [auth]): authorities, authorities_path, authz_base_url | R-AUTH-029 |
mount conflict, 503; backend initialization failed, 500; SFTP disconnection server configuration error: <cause>; see the server log | Combining roles, Types ([[roles.mounts]], [[backends]]): the application log gives the cause | R-AUTH-030, R-AUTH-031 |
username not usable as home directory | {username} ([[roles.mounts]]) | R-AUTH-032 |
session limit, session rejected: session limit exceeded for user <name> | The keys of [sftp] ([server]): max_sessions_per_user | R-SFTP-011 |
banned, 403 IP temporarily banned; SFTP: disconnection address banned, WARN address banned while this login was in flight — ... | The life of a ban, Lifting a ban ([sftp.ban], [api.ban], [admin.ban]) | R-BAN-008, R-BAN-008 |
rate limit, 429 rate limit exceeded | Rate limits ([sftp.rate_limit], [api.rate_limit]) | R-BAN-022 |
password checks saturated, 503 password checks saturated, retry later | Checking passwords ([auth]) | R-AUTH-011 |
password checks saturated for address | Checking passwords ([auth], [api.ban]): NAT, trusted_proxies | R-AUTH-012 |
shutting down | Shutdown | R-AUTH-015 |
invalid ticket, expired ticket, revoked ticket, 403 invalid download ticket, download ticket expired, download ticket revoked; 400 a download ticket is only redeemed by GET, ... is asked for with POST ?ticket, ... and an Authorization header are exclusive, ... only downloads a file, ?ticket and ?rename are exclusive | Downloading: the ticket, REST file API: ask for a ticket again | R-REST-009, R-REST-010 |
429 too many live download tickets | REST file API: 32 live tickets per user, wait for the oldest to expire | R-REST-009 |
range not satisfiable, 416 | REST file API: the Range starts after the end of the file, the client already has all of it | R-REST-006 |
login grace time exceeded | SSH: [sftp] ([sftp]): login_grace_secs | R-TIMEOUT-001 |
WARN failed to accept SFTP connections; retrying with backoff (descriptors exhausted, ulimit -n) | SFTP door | R-SFTP-002 |
SSH session ended: the peer offered no algorithm in common | An old client ([sftp.algorithms]) | R-SFTP-010 |
SSH request refused: ... (shell, exec, port forwarding); SSH_MSG_CHANNEL_FAILURE on a second sftp subsystem | only the sftp subsystem is served, once per connection: Doors | R-SFTP-012, R-SFTP-013 |
SSH_FX_FAILURE internal server error: this SFTP session is closed; ERROR a thread panicked: a defect in the server, ... (panic_payload, location, thread, backtrace with RUST_BACKTRACE=1) | a verb panicked; session_end reason=internal_error: to report | R-SFTP-019 |
File operations
reason of the operation lines, SFTP status and REST code.
reason | SFTP | REST | Where it is set | Rule |
|---|---|---|---|---|
acl | SSH_FX_PERMISSION_DENIED | 403 | Which entry decides ([[roles.mounts.acl]]) | R-ACL-003 |
acl subtree | SSH_FX_PERMISSION_DENIED | 403 | Rights ([[roles.mounts.acl]]): delete on the whole tree | R-DELETE-005 |
synthetic path | SSH_FX_PERMISSION_DENIED | 403 | Synthetic directories | R-ACL-005 |
rename across mounts | SSH_FX_OP_UNSUPPORTED | 422 | One mount or several | R-RENAME-010 |
invalid path | SSH_FX_PERMISSION_DENIED | 400 | control character, \, :, trailing dot, 8.3 short name: Client paths | R-UPLOAD-001 |
reserved name | SSH_FX_PERMISSION_DENIED | 403 | Two uploads to the same destination | R-RESERVE-010 |
rename into restricted | SSH_FX_PERMISSION_DENIED | 403 | Rights ([[roles.mounts.acl]]): write at the destination | R-RENAME-008 |
exists | SSH_FX_FAILURE | 409, 412 | existing destination | R-RENAME-002 |
is a directory, not a directory | SSH_FX_FAILURE | 409 | rm of a directory, rmdir of a file | R-DELETE-002 |
upload in progress; recursive delete refused: an upload is in progress beneath (WARN) | SSH_FX_PERMISSION_DENIED (curl: Permission denied (3)) | 409 | Two uploads to the same destination | R-RESERVE-001, R-RESERVE-012 |
quota exceeded | SSH_FX_FAILURE | 507 | Capping the size of a file ([[roles.mounts]]): max_file_mb | R-UPLOAD-007 |
session killed | SSH_FX_CONNECTION_LOST | - | session cut: Sessions | R-SFTP-020 |
unsupported | SSH_FX_OP_UNSUPPORTED | 501 | Atomic writes, Files and directories on S3 ([server.hidden_stores]): resume or append, send the whole file again | R-UPLOAD-009 |
upload in progress at CLOSE, client: upload not published: another upload took this file: <path> | SSH_FX_PERMISSION_DENIED | 409 | this upload’s lock was taken over by another instance before publication (refresh blocked): nothing is published, send the file again; Reservation across instances | R-RESERVE-008 |
taken over | SSH_FX_FAILURE on the old handle | 409 to the old upload | a stuck upload (suspended client) taken over by a retry from the same account: Two uploads to the same destination ([uploads]): takeover_idle_secs | R-RESERVE-017 |
upload idle timeout, upload below minimum rate | SSH_FX_FAILURE | 408 | Uploads: [uploads] ([uploads]) | R-UPLOAD-013 |
session ended: admin_kick | - | 503 the upload was cut by an administrator: nothing was written | a REST transfer cut from the console: Sessions | R-REST-013 |
session ended: <cause>, session ended | - | - | the client left; <cause> is the reason of the session_end | R-AUDIT-018 |
commit interrupted (result=unknown) | - | - | client left during publication: check the file | R-AUDIT-017 |
no roles | - | 403 | see no matching roles above | R-AUDIT-020 |
| - | SSH_FX_OP_UNSUPPORTED | - | READLINK, SYMLINK, posix-rename@openssh.com: not served | R-LIST-012 |
| - | - | 400 missing ?rename= query param; 405 (verb not served) | REST file API | R-REST-002 |
| - | SSH_FX_FAILURE | - | unknown handle or one of another kind: client bug | R-SFTP-017 |
Storage
A storage error carries its kind in reason (result=error); its
text is in the application log, at the same time.
reason, message | SFTP | REST | Where it is set | Rule |
|---|---|---|---|---|
not found | SSH_FX_NO_SUCH_FILE | 404 | Common keys on a local backend ([[roles.mounts]]): missing path; home_dir missing without create_home; on S3, a directory cannot be renamed | R-UPLOAD-015, R-S3-004 |
permission denied | SSH_FX_PERMISSION_DENIED | 403 | storage permissions; S3 without s3:ListBucket: Bucket permissions; WebHDFS (AuthorizationException, doAs not allowed): Prerequisites | R-UPLOAD-015, R-WEBHDFS-005 |
symlink escape | SSH_FX_PERMISSION_DENIED | 403 | The root and symbolic links ([[backends]] local): metric craftfilegate_local_symlink_refusals_total | R-LOCAL-004 |
already exists, directory not empty | SSH_FX_FAILURE | 409 | state of the storage | R-UPLOAD-015 |
storage error | SSH_FX_FAILURE backend error | 500 | the application log gives the cause | R-UPLOAD-015 |
not implemented | SSH_FX_OP_UNSUPPORTED | 501 | an operation this backend does not have: Types | R-UPLOAD-015 |
root <path> cannot be canonicalized and opened as a directory: <error> | - | - | The root and symbolic links ([[backends]] local) | R-LOCAL-003 |
no host_key_fingerprint: the upstream's host key would not be checked, ...; SFTP proxy: the upstream's host key does not match host_key_fingerprint (ERROR) | - | - | The upstream’s host key ([[backends]] sftp): ssh-keyscan -p <port> <host> | ssh-keygen -lf - | R-PROXY-002, R-PROXY-004 |
SFTP proxy: authentication failed: the upstream accepts only ssh-rsa (SHA-1) ... | - | - | Keys ([[backends]] sftp): ed25519 or ECDSA key | R-PROXY-005 |
WARN this store refused the conditional CopyObject with a 400 this code does not treat as "precondition not implemented": ... (every rename fails); could not list the multipart uploads under this backend's prefix, ..., listing the parts of a multipart upload was refused, ..., could not read the S3 service's clock (no usable Date header on its response) and ..., this S3 backend has no prefix, so its abandoned multipart uploads are never swept: ... | - | - | Bucket permissions, Multipart uploads, Conditional writes | R-S3-010, R-S3-013 |
SFTP proxy connect: ...; WARN accept_any_host_key = true on this SFTP proxy backend: ..., sftp proxy: the upstream cannot replace a file atomically (posix-rename@openssh.com), ..., sftp proxy: the second SFTP channel used for posix-rename@openssh.com could not be opened (...), ... | - | - | The upstream’s host key, Uploads | R-PROXY-001, R-PROXY-003, R-PROXY-007 |
ERROR sftp proxy: the destination was removed to publish an upload and the rename ... which is kept, sftp proxy: an upload was interrupted after the removal of its destination ... which is kept (fields temp, destination) | - | - | Uploads: the in-flight file is the only copy, rename it by hand | R-PROXY-008 |
the conditional writes are not proven on this S3 store: <why> (WARN), refusal with cross_instance_reservation; INFO S3 upload precondition self-test verdict = ignored, not implemented, if-match ignored, delete refused, inconclusive | - | - | Conditional writes ([[backends]]): cross_instance_reservation | R-S3-007, R-RESERVE-015 |
the SFTP upstream refuses the name of the upload lock files, ..., could not refresh the lock file of an upload in progress: ..., this upload's lock was taken over by another instance: it is not published (WARN) | - | - | Reservation across instances ([[backends]]): lock_prefix | R-RESERVE-005, R-RESERVE-007, R-RESERVE-008 |
backend "<name>" (webhdfs): Knox refused the service account on GETFILESTATUS of the root ...; session refused: a username that cannot be a doAs; WARN this WebHDFS backend's url is plain http ..., this WebHDFS backend's gateway could not be checked at startup ..., this WebHDFS answers no LISTSTATUS_BATCH ... | - | - | Prerequisites, The mount on WebHDFS ([[backends]] webhdfs): auth, url, ca_bundle | R-WEBHDFS-004, R-WEBHDFS-005, R-WEBHDFS-001, R-WEBHDFS-010 |
case_insensitive = true on a WebHDFS backend: ... | - | - | The mount on WebHDFS ([[backends]]) | R-WEBHDFS-014 |
the storage's clock differs from this server's by more than 30 s (WARN); stale in-flight files on this SFTP upstream are never collected (age_check = false): ..., stale in-flight files of this backend are aged on this server's clock (age_check = false): ..., abandoned multipart uploads of this S3 backend are aged on this server's clock (age_check = false): ..., could not read the storage's clock with a probe file in this directory (age_check is on), ..., this storage offers no lock to tell an in-flight upload from the leftovers of one killed with the server ... | - | - | Leftovers of an interrupted transfer ([uploads.stale_partials]): NTP of the storage | R-HIDDEN-016, R-HIDDEN-011, R-HIDDEN-013 |
this filesystem does not support RENAME_NOREPLACE ... (WARN) | - | - | Performance and limits: NFS, FUSE | R-LOCAL-011 |
backend unavailable: its storage did not answer the availability probe (WARN, backend, backend_type, failures, reason); backend available again: ... (WARN, down_secs) | - | - | Availability ([server.backend_probe]): reason says what the visit found (missing root, bucket refused, upstream unreachable, different host key, timeout) | R-AVAIL-002 |
reason SFTP proxy: authentication failed — probe paused until reload | - | - | SFTP proxy: the service account or its password; the probe reconnects only at the next reload of the roles, so the upstream does not ban the gateway | R-AVAIL-001 |
Admin console and API
Code and detail | Where it is set | Rule |
|---|---|---|
401 missing or invalid authorization header, invalid token | Authenticating ([admin], [auth.jwt]) | R-ADMIN-005 |
401 invalid credentials on POST /admin/login (counts for [admin.ban]; an account without an admin role: admin sign-in refused: the password is right, ... at WARN), 503 password checks saturated, retry later, 400 expected a JSON body {"username": ..., "password": ...}, 415 expected Content-Type: application/json (a client that does not send Content-Type: application/json), 403 sign-in from another site refused (Sec-Fetch-Site other than same-origin or none: a page from another site; neither counts for the ban), 400 username longer than 256 bytes; 401 revoked token (account removed, password changed; or revoked, see below), 401 session too old: sign in again, 400 only a session token is renewed: ... | Signing in ([admin.session], users_file): sign in again | R-ADMIN-022 |
404 password sign-in is not available on this door on POST /admin/login; [admin.session] sets a session key or a revocation store, but no account signs in to the console: ... (WARN); console: only the token field, no password form (GET /admin/login says {"password": false}) | Console sessions ([auth.methods], [[admin.roles]]): local passwords and admin roles | R-ADMIN-022, R-ADMIN-018 |
401 revoked token after POST /admin/revocations or POST /admin/logout: sign in again; lift local_only and WARN the session tokens were revoked on this instance only: ...; admin session revocations are kept in memory beside a shared session key: ..., [admin.session] sets revocation keys this store does not read, shared admin session revocation records refused ..., ... ahead of this clock ... (WARN); ERROR the shared admin session revocations cannot be read: ...; refusal to start ... the revocations cannot be read ..., admin.session.backend = ..., admin.session.reread_interval_secs must be between 1 and 30 ..., admin.session.revocation_configmap_name is empty ..., admin.session.persist_file <path> is also the persist_file of a ban list ...; 400 expected a JSON body {"username": ...}, username blank or longer than 256 bytes, 404 this door issues no session token: there is nothing to revoke | Revoking, signing out, Sharing revocations ([admin.session]): persist_file or backend, NTP | R-ADMIN-023 |
POST /admin/grants: 400 unknown role, until in the past, invalid period, duration over max; 403 role not grantable; 409 role already held, mount conflict, already granted, too many grants; 403 the grant permission is required; 404 this door grants no temporary access (listener without [admin]); DELETE: 404 grant not found, 409 grant ended, grant ended: it is already <state>; lift local_only; refusal to start admin.grants.* ..., [cluster] is set and the temporary access grants are kept in memory: ..., ... the grants cannot be read ...; WARN [admin.grants] is set, but no credential holds the grant permission: ...; SFTP session cut temporary access expired / temporary access revoked, REST upload 503 the upload was cut: the temporary access it used ended; nothing was written; WARN a temporary access grant names a role this server does not define: ignored, ... names an admin role: ignored; ERROR the shared temporary access grants cannot be read: ... | Reference [admin.grants]: max_duration_secs, persist_file, backend | R-GRANT-002, R-GRANT-006, R-GRANT-003, R-GRANT-005, R-GRANT-009, R-GRANT-010 |
console: Wrong username or password. (a wrong password, an unknown name or an account without an admin role: a single message); Your session was revoked: sign in again., Your session expired: sign in again., Your session reached its maximum duration: sign in again., Token rejected: sign in again.; the page asks to sign in again after a reload or in a new tab | In the console ([admin.session]): sign in again; the token lives only in the page | R-ADMIN-018 |
blank or unstyled console behind a reverse proxy, Refused to ... / violates the following Content Security Policy directive in the browser console | The console: the proxy’s CSP allows at least the page’s | R-ADMIN-018 |
403 no matching admin roles; 403 the <permission> permission is required | Authenticating ([[admin.roles]]) | R-ADMIN-006, R-ADMIN-007 |
404 session not found: <id>, 400 invalid session ID format | Sessions | R-ADMIN-010 |
404 no log files: [log] dir is not set, 404 unknown log source: expected app or audit, 500 the log file could not be read, 403 the log file is a symbolic link, which the log viewer does not follow; 400 lines: not a number, level: unknown level, q: longer than 256 bytes, field: ...; 429 too many log reads at once: retry in a moment | Log files, Logs and streams ([log]) | R-ADMIN-013 |
429 too many admin streams open: close one or retry later | Logs and streams | R-ADMIN-014 |
unban 409 lift=overruled (a ban later than the request applies), 200 lift=local_only (shared state not written), 404 not banned, no ban manager, 400 unknown protocol (sftp, api or admin) | Lifting a ban, Sharing bans between instances ([sftp.ban], [api.ban], [admin.ban]) | R-BAN-011 |
401 (404 with the file explorer, [api.ui]) on /metrics, /admin, /health, /livez or /readyz called on admin.listen; 404 on <prefix>: [api] enabled missing, or request on control_listen | One door or two ([admin], [api]): with control_listen or probes_listen, each route has its port | R-METRICS-001, R-REST-001 |
404 or 401 on /api/docs, /api/openapi.json | REST file API ([api]): openapi = true | R-REST-012 |
port unreachable from outside a container, connection refused; container unhealthy (craft-file-gate healthcheck failing) | Docker, The binary’s tools ([admin]): listen = "0.0.0.0:...", not 127.0.0.1; the address that healthcheck queries | R-ADMIN-001, R-ADMIN-015 |
the admin door shares its listener with the file API ... (WARN) | One door or two ([admin]): set control_listen | R-ADMIN-002 |
admin TLS certificate expires soon, has expired, admin TLS certificate expiry could not be read | What you will see ([admin.tls]): renew; an expired certificate is served anyway | R-ADMIN-004 |
Kubernetes
| Symptom or message | Where it is set | Rule |
|---|---|---|
probes failing, connection refused | Kubernetes ([sftp], [admin]): listen = "0.0.0.0:..."; probes blocked by the NetworkPolicy depending on the CNI: the node CIDR in networkPolicy.controlFrom | R-ADMIN-001 |
/readyz 503, data_runtime=unresponsive (runtime saturated) or sftp=not_accepting (SFTP port closed, shutdown in progress) | Probes | R-K8S-002 |
chart rendering fails: networkPolicy.controlFrom is empty ..., config.existingConfigMap and config.inline are both set, no configuration: ..., adminRevocations.backend = ...: expected memory, file, configmap or empty, replicaCount > 1 with clusterSecret.enabled = false ..., replicaCount > 1 with a shared console session key and adminRevocations kept in memory ..., cluster is on (replicaCount > 1, or cluster.enabled) and clusterSecret.enabled is false ..., cluster.unreadyWhen names <value>: ... | The chart, The pods’ shared Secret | R-K8S-005 |
backend = "configmap" refused: ConfigMap unreadable for 30 s (message naming <namespace>/<name> and the Role); WARN at each re-read of the bans, propagation in 30 s (watch right missing) | Sharing through a ConfigMap ([sftp.ban], [api.ban], [admin.ban]): Role and RoleBinding | R-BAN-017 |
cannot read or fill the entry admin-session.key of Secret <ns>/<name> ..., shared Secret not readable or writable yet; ... (WARN) | The pods’ shared Secret ([admin.session]): the Secret created by the chart, the Role (get, update on that name) | R-ADMIN-022 |
cannot read the admin session revocations in ConfigMap <ns>/<name> ..., admin session revocations not readable yet; ... (WARN) | Sharing revocations: the chart’s ConfigMap and Role (adminRevocations.backend: configmap, get, update, patch on that name) | R-ADMIN-023 |
cannot read or fill the entry tls.key of Secret <ns>/<name> ..., the entries tls.crt and tls.key of Secret <name> do not go together ..., ... the key does not belong to the certificate; WARN cluster peers name does not resolve: its last addresses are kept; WARN cluster peer unreachable (only from a peer that has already answered; before that, at DEBUG), cluster peer presents another certificate than this instance ..., cluster state cut to fit: ..., cluster peer state holds absurd records ..., cluster peer's failure series are all past their window ... (clocks to synchronize, NTP) | The shared certificate: the chart’s Secret and Role, a rotation in progress; the cluster port in the NetworkPolicy | R-CLUSTER-002, R-CLUSTER-006, R-CLUSTER-005, R-CLUSTER-003 |
Instances tab, peers of a scope=cluster list, relayed kick or lift: unauthorized on <instance>, forbidden on <instance>; WARN audit: relayed credential refused ... on the peer; 400 unknown scope: ... (scope other than local or cluster) | Cluster: the same static token, the same session key, the same [auth.jwt] and the same [[admin.roles]] everywhere | R-CLUSTER-009, R-CLUSTER-010, R-CLUSTER-008 |
backend = "configmap" refused by a binary without the k8s feature | The binary’s features | R-BAN-018 |
failed to create the file watcher, falling back to re-reading ... | inotify on a shared node ([reload]) | R-RELOAD-010 |
pod OOMKilled during a burst of connections | Memory, Checking passwords ([auth]) | R-AUTH-014 |
“host key changed” from one pod to another; WARN host key not found, generated a new one (sftp.generate_host_key) | SSH host key ([sftp]) | R-SFTP-004 |
Logs, metrics and telemetry
| Message | Where it is set | Rule |
|---|---|---|
craft-file-gate: <n> log line(s) could not be written ... (stderr) | An incomplete trail | R-AUDIT-031 |
audit target disabled by RUST_LOG, audit successes disabled by RUST_LOG | RUST_LOG (Environment variables): RUST_LOG=warn,audit=info | R-AUDIT-002 |
process resource sampling failed; the process_* series are absent from /metrics | Metrics | R-METRICS-010 |
OTLP collector unreachable — telemetry spans will be dropped until recovery | Exports, retries and shutdown ([telemetry]) | R-TELEMETRY-010 |
failed to create the OTLP <signal> exporter, ... ([telemetry] on_exporter_error = "warn") (ERROR) | [telemetry] ([telemetry]) | R-TELEMETRY-003 |
sessions were still ending when the grace period ran out: ... | Tuning the orchestrator ([server]): shutdown_grace_period_secs | R-SHUTDOWN-007 |
ban file was still being written when the shutdown stopped waiting; the shutdown stopped waiting for the removal of upload lock files; ... | Shutdown | R-BAN-019, R-SHUTDOWN-008 |
Reference: configuration
The whole configuration fits in one TOML file, passed with
craft-file-gate --config <file>. Users, roles and backends can also live
in users_file and roles_file, next to it.
| Page | Sections |
|---|---|
[server], [cluster] | [server], [server.hidden_stores], [cluster] |
[sftp] and the SSH door | [sftp], [sftp.algorithms], [sftp.ban], [sftp.rate_limit] |
[auth], users and roles | [auth], [auth.methods], [auth.jwt], [[users]], [[roles]], [[roles.mounts]], [[roles.mounts.acl]] |
[[backends]] | common keys, then local, sftp, s3, webhdfs |
[admin] and [api] | [admin], [[admin.roles]], [admin.tls], [admin.ban], [admin.metrics_thresholds], [api], [api.ban], [api.rate_limit], [api.ui] |
[log], [telemetry] and the rest | [log], [telemetry], [uploads], [uploads.stale_partials], [tcp_keepalive], [reload], [security] |
Only [auth] is required, with at least one role, plus at least one door:
[sftp], or [api] with [admin]. With no door configured, the server
refuses to start.
craft-file-gate config explain <key> gives each key its type, its default and
its bounds (Checking a configuration).
Reading these tables
| Column | Meaning |
|---|---|
| Key | the full path; [] marks an entry of an array of tables: roles[].mounts[].home_dir is the home_dir key of a [[roles.mounts]]; ⟳: reloaded at runtime |
| Type | string, integer, boolean, list, table, array of tables, path, address (ip:port), URL |
| Default | the value when the key is absent; (required): its absence refuses startup; -: no value of its own |
| Effect | what the key sets, and its bounds |
A key without a mark waits for a restart. An open session keeps its roles and its mounts until it ends. See Hot reload.
Common rules
| Rule | Effect |
|---|---|
| relative path | relative to the directory of the configuration file, not to the current directory |
roles.toml, users.toml next to the configuration | used if not declared; the startup log says which |
| environment variable | replaces the value from the file: see Environment variables |
| unknown key, out-of-bounds value, key with no effect | Troubleshooting |
Reference: [server] and [cluster]
⟳: reloaded at runtime; no mark: taken at restart. See Hot reload.
What all doors share, and the channel between instances. [server] can be missing: each key has its default.
[server]
See Shutdown, Kubernetes.
| Key | Type | Default | Effect |
|---|---|---|---|
server | table | - | the settings common to all doors |
server.shutdown_grace_period_secs | integer | 30 | seconds left to the work in progress of each door at shutdown, one deadline for all; at least 1 |
server.max_sessions_per_user | integer | unlimited | concurrent sessions per user, on a door with sessions (SFTP); with [cluster], those of each reachable instance add up |
server.probes_listen | address | absent: probes on [admin] | CRAFT_FILE_GATE_PROBES_LISTEN replaces it; a separate port for /livez, /readyz, /health and /metrics, served only there, over HTTP, without authentication; without [admin], the only HTTP listener |
[server.hidden_stores]
See Atomic writes.
| Key | Type | Default | Effect |
|---|---|---|---|
server.hidden_stores | table | - | atomic writes, for every backend and every mount that says nothing |
server.hidden_stores.enabled | boolean | false | write to an in-transfer file, published by a rename at the end |
server.hidden_stores.prefix | string | .in. | the start of that file’s name |
server.hidden_stores.extension | string | . | the end of that file’s name |
[server.backend_probe]
See Backends.
| Key | Type | Default | Effect |
|---|---|---|---|
server.backend_probe | table | - | the availability probe of each backend, in the background, never on the path of an operation |
server.backend_probe.enabled | boolean | true | false: no probe, each backend stays unknown |
server.backend_probe.interval_secs | integer | 15 | seconds between two visits to a backend, from 5 to 3600 |
server.backend_probe.timeout_secs | integer | 5 | maximum duration of a visit, from 1 to 60, below interval_secs (otherwise startup is refused) |
server.backend_probe.failures_before_down | integer | 2 | failed visits in a row before down, from 1 to 10; one successful visit is enough to go back up. Worst-case detection: interval_secs × failures_before_down + timeout_secs, 35 s by default |
[cluster]
See Cluster.
| Key | Type | Default | Effect |
|---|---|---|---|
cluster | table | absent: each instance is alone | the channel between the instances of a deployment; CRAFT_FILE_GATE_CLUSTER_LISTEN creates it |
cluster.listen | address | - | the channel’s own port, in mutual TLS 1.3; CRAFT_FILE_GATE_CLUSTER_LISTEN replaces it |
cluster.peers | string or list | - | dns:<name>:<port>: each address of the name, re-resolved every 5 s (a headless Service yields each ready pod, Docker Compose each replica of the service); or <host>:<port> entries. Each one is called every second; the one that answers the instance’s own identifier (<pod_name>/<boot_id>) is itself; CRAFT_FILE_GATE_CLUSTER_PEERS replaces it |
cluster.cert_file | path | - | the shared certificate; when it and key_file are missing, generated (ECDSA P-256, 10 years) by the first instance, read by the others on a shared volume |
cluster.key_file | path | - | its key, created with 0600 |
cluster.secret_name | string | - | or the shared Kubernetes Secret, entries tls.crt and tls.key (feature k8s); CRAFT_FILE_GATE_CLUSTER_SECRET_NAME replaces it |
cluster.peer_timeout_ms | integer | 2000 | timeout of a call to a peer, from 100 to 10000 |
cluster.min_peers | integer | 1 | below this number of reachable peers for isolated_after_secs, the instance is isolated; 0: never; from 0 to 1000; CRAFT_FILE_GATE_CLUSTER_MIN_PEERS replaces it |
cluster.unready_when | list | []: alert only | the detectors that withdraw the instance from service (/readyz not ready), ORed: storage_alone, isolated, isolated_and_storage_down; CRAFT_FILE_GATE_CLUSTER_UNREADY_WHEN (comma-separated) replaces it |
cluster.isolated_after_secs | integer | 30 | time below min_peers before being isolated; immediate return; from 5 to 3600 |
cluster.unready_after_secs | integer | 60 | duration of a chosen detector before withdrawal, from 10 to 3600, at least 2 × server.backend_probe.interval_secs |
cluster.ready_after_secs | integer | 30 | duration without a chosen detector before returning, from 5 to 3600, at least server.backend_probe.interval_secs |
Reference: [sftp] and the SSH door
⟳: reloaded at runtime; no mark: taken at restart. See Hot reload.
[sftp]
See The SFTP door, Timeouts.
| Key | Type | Default | Effect |
|---|---|---|---|
sftp | table | absent: SFTP door off | the SFTP door (feature door-sftp) |
sftp.listen | address | (required) | listen address and port; CRAFT_FILE_GATE_LISTEN replaces it |
sftp.host_keys | list | (required) | the host keys, existing files: a path, or a { path, algorithms } table |
sftp.host_keys[].path | path | (required) | the private key file, table form |
sftp.host_keys[].algorithms | list | depends on the key | the signature algorithms advertised for this key, [sftp.algorithms] syntax |
sftp.generate_host_key | string | absent | "ed25519" or "ecdsa-p256": creates the key of a missing host_keys file; single instance only |
sftp.login_grace_secs | integer | 120 | seconds to authenticate, connection closed beyond that; 0 disables |
sftp.server_id | string | CraftFileGate_<version> | the SSH-2.0-<server_id> banner |
sftp.inactivity_timeout_secs | integer | 600 | a connection with no packet is closed after this timeout; 0 never; at most 86400 |
sftp.keepalive_interval_secs | integer | 0 | SSH keepalive after this client silence; 0 none; at most 86400 |
sftp.keepalive_max | integer | 3 | unanswered keepalives before closing; 1 to 100 |
[sftp.algorithms]
| Key | Type | Default | Effect |
|---|---|---|---|
sftp.algorithms | table | - | the SSH algorithms: +name adds to the default, -name removes, a list without prefix replaces |
sftp.algorithms.kex | list | mlkem768x25519-sha256, curve25519-sha256, curve25519-sha256@libssh.org, ecdh-sha2-nistp256, ecdh-sha2-nistp384, ecdh-sha2-nistp521, diffie-hellman-group16-sha512, diffie-hellman-group14-sha256 | key exchange |
sftp.algorithms.ciphers | list | chacha20-poly1305@openssh.com, aes256-gcm@openssh.com, aes128-gcm@openssh.com, aes256-ctr, aes192-ctr, aes128-ctr | ciphers |
sftp.algorithms.macs | list | hmac-sha2-512-etm@openssh.com, hmac-sha2-256-etm@openssh.com, hmac-sha2-512, hmac-sha2-256 | MAC |
sftp.algorithms.host_key | list | ssh-ed25519, ecdsa-sha2-nistp256, ecdsa-sha2-nistp384, ecdsa-sha2-nistp521, rsa-sha2-512, rsa-sha2-256 | host key signatures |
sftp.algorithms.user_key | list | ssh-ed25519, ecdsa-sha2-nistp256, ecdsa-sha2-nistp384, ecdsa-sha2-nistp521, sk-ssh-ed25519@openssh.com, sk-ecdsa-sha2-nistp256@openssh.com, rsa-sha2-512, rsa-sha2-256 | signatures allowed for a user key; ssh-rsa (SHA-1) not in the default |
[sftp.ban]
See Bans.
| Key | Type | Default | Effect |
|---|---|---|---|
sftp.ban | table | absent: no ban | ban of addresses after authentication failures on the SFTP door |
sftp.ban.max_failures | integer | 5 | failures in the window before the ban; at least 1 |
sftp.ban.ban_duration_secs | integer | 600 | duration of a ban in seconds, from the verdict |
sftp.ban.window_secs | integer | 300 | failure counting window in seconds, fixed, opened by the first failure of a series |
sftp.ban.whitelist_ips | list | [] | addresses and CIDR networks, IPv4 or IPv6, never banned by this instance |
sftp.ban.trusted_proxies | list | [] | no effect on SSH |
sftp.ban.persist_file | path | absent: in memory | file where bans are kept and shared between instances |
sftp.ban.backend | string | "file" | "file", or "configmap" to share bans between Kubernetes pods |
sftp.ban.ban_configmap_name | string | craft-file-gate-bans | the ConfigMap of bans, with backend = "configmap" |
sftp.ban.reread_interval_secs | integer | 5 | re-read of the shared persist_file in seconds, on top of file watching; 1 to 30 |
[sftp.rate_limit]
| Key | Type | Default | Effect |
|---|---|---|---|
sftp.rate_limit | table | absent: no limit | token bucket per address, when a connection is accepted |
sftp.rate_limit.connections_per_minute | integer | (required) | sustained connection rate per address; at least 1 |
sftp.rate_limit.burst | integer | connections_per_minute | connections accepted in a row; at least 1 |
Reference: [auth], users and roles
⟳: reloaded at runtime; no mark: taken at restart. See Hot reload.
[auth]
See Authentication, Password hashing.
| Key | Type | Default | Effect |
|---|---|---|---|
auth | table | (required) | authentication and the role sources |
auth.jwt_sentinel_username | string | jwt | the SSH name that requests JWT authentication (the token as password) |
auth.timeout_secs | integer | 5 | timeout of a call to the authorization service |
auth.authz_base_url | URL | absent | service that translates authorities into role names (POST /authz/resolve); without it, only authorities that are role names count |
auth.roles_file | path | roles.toml next to it, if it exists | file of [[roles]] and [[backends]]; its content is reloaded at runtime |
auth.users_file | path | users.toml next to it, if it exists | file of [[users]]; its content is reloaded at runtime |
auth.hash_workers | integer | the CPUs seen (cgroup quota included), at most 2 | threads that verify passwords, concurrently; each holds the argon2 m of its hash; 1 to 1024; CRAFT_FILE_GATE_HASH_WORKERS |
auth.hash_queue | integer | 1024 | verifications that can wait for a thread; 1 to 65536; CRAFT_FILE_GATE_HASH_QUEUE |
auth.hash_per_address | integer | 4 | verifications in flight or queued for one address; 1 to 66560; no cap for an address in the whitelist_ips of the door’s ban list |
auth.users ⟳ | array of tables | - | [[auth.users]], like [[users]]; the root wins |
auth.roles ⟳ | array of tables | - | [[auth.roles]], like [[roles]]; the root wins |
auth.backends ⟳ | array of tables | - | [[auth.backends]], like [[backends]]; the root wins |
[auth.methods]
| Key | Type | Default | Effect |
|---|---|---|---|
auth.methods | table | - | the active identity sources |
auth.methods.jwt | table | - | the JWT method |
auth.methods.jwt.enabled | boolean | true | accept JWTs (SFTP under the sentinel name, REST as Bearer); requires [auth.jwt] |
auth.methods.local | table | - | the local user store |
auth.methods.local.enabled | boolean | false | enable the local store ([[users]] or users_file) |
auth.methods.local.password | boolean | true | accept a password as proof |
auth.methods.local.pubkey | boolean | true | accept an SSH public key as proof |
auth.methods.local.allow_sha512_crypt | boolean | false | also accept sha512-crypt hashes, $6$ (migration) |
auth.methods.local.allow_bcrypt | boolean | false | also accept bcrypt hashes, $2a$, $2b$, $2y$ (migration) |
auth.methods.local.users_file | path | absent | users file, like auth.users_file |
auth.methods.local.max_argon2_memory_kib | integer | absent: nothing is checked | memory budget of one verification, in KiB; a hash that exceeds it is named at load time; at least 1 |
[auth.jwt]
| Key | Type | Default | Effect |
|---|---|---|---|
auth.jwt | table | absent: no JWT verified | JWT verification: a key source, and the claims |
auth.jwt.secret | string | absent | HMAC secret (HS256/384/512); CRAFT_FILE_GATE_JWT_SECRET or _FILE |
auth.jwt.public_key_file | path | absent | PEM public key (RS*, ES*) |
auth.jwt.jwks_url | URL | absent | JWKS endpoint of the identity provider; exclusive with secret and public_key_file |
auth.jwt.jwks_refresh_interval_secs | integer | 3600 | JWKS refresh period; at least 1 |
auth.jwt.algorithm | string | HS256 | HS256, HS384, HS512, RS256, RS384, RS512, ES256 or ES384 |
auth.jwt.username_path | string | /sub | JSON Pointer to the user name in the token |
auth.jwt.authorities_path | string | /groups | JSON Pointer to the user’s authorities (roles) |
auth.jwt.issuer | string | absent | if set, iss is required and compared |
auth.jwt.audience | string | absent | if set, aud is required and compared |
[[users]]
| Key | Type | Default | Effect |
|---|---|---|---|
users ⟳ | array of tables | - | the local users; also in users_file |
users[].username ⟳ | string | (required) | the name; unique |
users[].password_hash ⟳ | string | (required) | argon2 hash (craft-file-gate hash-password); never in clear text |
users[].authorized_keys ⟳ | list | [] | SSH public keys, one authorized_keys line each |
users[].authorities ⟳ | list | [] | the names of the user’s roles |
[[roles]]
| Key | Type | Default | Effect |
|---|---|---|---|
roles ⟳ | array of tables | (required), here or in roles_file | the roles; also in roles_file |
roles[].name ⟳ | string | (required) | the role name, unique; it is what authorities cite |
roles[].user_key_algorithms ⟳ | list | [] | SSH key signature algorithms additionally allowed to the local users of this role, by name (see The SFTP door) |
roles[].mounts ⟳ | array of tables | (required) | the role’s mounts, [[roles.mounts]]; at least one |
[[roles.mounts]]
| Key | Type | Default | Effect |
|---|---|---|---|
roles[].mounts[].backend ⟳ | string | (required) | the mounted backend, by the name of a [[backends]] |
roles[].mounts[].mount_path ⟳ | path | / | where the mount appears to the user: absolute, without ., .., // or a trailing / |
roles[].mounts[].home_dir ⟳ | path | / | where the mount starts on the storage; {username} there becomes the user’s name |
roles[].mounts[].create_home ⟳ | boolean | false | create home_dir at login if missing; local backend only, see Local |
roles[].mounts[].max_file_mb ⟳ | integer | absent: no cap | maximum size of an uploaded file, in MB (1,048,576 bytes); 0: no upload |
roles[].mounts[].acl ⟳ | array of tables | []: everything refused | the mount’s rights, [[roles.mounts.acl]], see ACL |
roles[].mounts[].hidden_stores ⟳ | table | the backend’s | atomic writes of this mount, key by key (see Atomic writes) |
roles[].mounts[].hidden_stores.enabled ⟳ | boolean | the backend’s | see server.hidden_stores.enabled |
roles[].mounts[].hidden_stores.prefix ⟳ | string | the backend’s | see server.hidden_stores.prefix |
roles[].mounts[].hidden_stores.extension ⟳ | string | the backend’s | see server.hidden_stores.extension |
[[roles.mounts.acl]]
| Key | Type | Default | Effect |
|---|---|---|---|
roles[].mounts[].acl[].path ⟳ | path | (required) | governed path, relative to the mount |
roles[].mounts[].acl[].rights ⟳ | list | (required) | among read, write, list, delete, rename |
roles[].mounts[].acl[].recursive ⟳ | boolean | false | true: the entry also governs everything under path |
Reference: [[backends]]
⟳: reloaded at runtime; no mark: taken at restart. See Hot reload.
[[backends]]: all types
See Backends.
| Key | Type | Default | Effect |
|---|---|---|---|
backends ⟳ | array of tables | - | the storages; also in roles_file |
backends[].name ⟳ | string | (required) | the name that mounts cite; unique |
backends[].type ⟳ | string | (required) | local, sftp, s3 or webhdfs, among those built into the binary |
backends[].stale_partials ⟳ | table | that of [uploads.stale_partials] | sweep of the leftovers of interrupted uploads on this backend |
backends[].stale_partials.age_check ⟳ | boolean | that of [uploads.stale_partials] | see uploads.stale_partials.age_check |
backends[].stale_partials.grace_secs ⟳ | integer | that of [uploads.stale_partials] | see uploads.stale_partials.grace_secs |
backends[].hidden_stores ⟳ | table | that of [server.hidden_stores] | atomic writes of this backend, key by key; not applicable on S3 and WebHDFS |
backends[].hidden_stores.enabled ⟳ | boolean | that of [server.hidden_stores] | see server.hidden_stores.enabled |
backends[].hidden_stores.prefix ⟳ | string | that of [server.hidden_stores] | see server.hidden_stores.prefix |
backends[].hidden_stores.extension ⟳ | string | that of [server.hidden_stores] | see server.hidden_stores.extension |
backends[].refuse_upload_over_directory ⟳ | boolean | false | S3: refuse an upload to a key that is also a directory; not applicable elsewhere, where an upload never replaces a directory |
backends[].cross_instance_reservation ⟳ | boolean | true (false for local on Windows) | put a lock on the storage during an upload, for the other instances; false suits a single instance |
backends[].lock_prefix ⟳ | string | .craftfilegate-upload. | the start of the name of these locks; 8 to 64 bytes, without / |
backends[].case_insensitive ⟳ | boolean | probed (local), false (elsewhere) | compare ACL paths ignoring case and Unicode normalization; absent on WebHDFS |
[[backends]] type = “local”
See Local.
| Key | Type | Default | Effect |
|---|---|---|---|
backends[].root ⟳ | path | (required) | local: the root directory |
backends[].follow_symlinks ⟳ | boolean | false | local: follow symbolic links that stay under root + the mount’s home_dir |
[[backends]] type = “sftp”
See SFTP proxy.
| Key | Type | Default | Effect |
|---|---|---|---|
backends[].host ⟳ | string | (required) | SFTP proxy: the upstream |
backends[].port ⟳ | integer | 22 | SFTP proxy: its port; 1 to 65535 |
backends[].host_key_fingerprint ⟳ | string | (required, unless accept_any_host_key) | SFTP proxy: the fingerprint of the upstream host key, SHA256:<base64> (as ssh-keygen -lf) or SHA512:<base64> |
backends[].accept_any_host_key ⟳ | boolean | false | SFTP proxy: true without a fingerprint, any host key is accepted; for a throwaway upstream (tests, mock-ups) only |
backends[].auth ⟳ | table | (required) | SFTP proxy, WebHDFS: the service account |
backends[].auth.type ⟳ | string | (required) | SFTP proxy: password or private_key; WebHDFS: basic |
backends[].auth.username ⟳ | string | (required) | the service account; WebHDFS: without : or control characters |
backends[].auth.password ⟳ | string | - | SFTP proxy, type = "password": its password |
backends[].auth.private_key_pem ⟳ | string | - | SFTP proxy, type = "private_key": its PEM private key |
[[backends]] type = “s3”
See S3 and compatibles.
| Key | Type | Default | Effect |
|---|---|---|---|
backends[].bucket ⟳ | string | (required) | S3: the bucket |
backends[].region ⟳ | string | (required) | S3: the region |
backends[].prefix ⟳ | string | "" | S3: the start of every key |
backends[].endpoint_url ⟳ | URL | AWS S3 | S3: the endpoint of a compatible service (MinIO, Garage…) |
backends[].credentials ⟳ | table | (required) | S3: { type = "iam_role" } or { type = "static", ... } |
backends[].credentials.type ⟳ | string | (required) | S3: static (a key pair) or iam_role (the environment’s chain) |
backends[].credentials.access_key_id ⟳ | string | - | S3, static: the access key |
backends[].credentials.secret_access_key ⟳ | string | - | S3, static: its secret |
[[backends]] type = “webhdfs”
See WebHDFS (Knox). auth, auth.type and auth.username: SFTP proxy section above.
| Key | Type | Default | Effect |
|---|---|---|---|
backends[].url ⟳ | URL | (required) | WebHDFS: the base of the Knox gateway, https://<knox>:<port>/gateway/<topology>; the backend appends /webhdfs/v1; without credentials, query or fragment |
backends[].auth.password_file ⟳ | path | (required) | WebHDFS: file holding the service account password, re-read at each reload of the file that defines the backend; a trust anchor |
backends[].ca_bundle ⟳ | path | system store | WebHDFS: the PEM roots that authenticate url, alone; a trust anchor |
backends[].impersonate ⟳ | boolean | true | WebHDFS: doAs=<user> on each request; false: everything goes out under the service account |
backends[].read_ahead_bytes ⟳ | integer | 4194304 | WebHDFS: size of a read range; 65536 to 67108864 |
backends[].timeout_secs ⟳ | integer | 30 | WebHDFS: timeout of a request in seconds, body included; at least 1 |
Reference: [admin] and [api]
⟳: reloaded at runtime; without the mark: taken at restart. See Hot reload.
[admin]
| Key | Type | Default | Effect |
|---|---|---|---|
admin | table | absent: no HTTP door | the listener of the admin API, the file API and the file explorer |
admin.listen | address | (required) | listen address; 0.0.0.0 in a container; CRAFT_FILE_GATE_ADMIN_LISTEN |
admin.control_listen | address | absent | a separate door for /admin, the console, /metrics and the probes (except with server.probes_listen); CRAFT_FILE_GATE_ADMIN_CONTROL_LISTEN |
admin.bearer_token | string | absent | token with every permission; it or a role from [[admin.roles]] is required; CRAFT_FILE_GATE_ADMIN_BEARER_TOKEN, _FILE |
admin.allow_static_token | boolean | true | false: no static token, from any source |
admin.header_read_timeout_secs | integer | 10 | seconds to send the headers of a request; 1 to 300; CRAFT_FILE_GATE_ADMIN_HEADER_READ_TIMEOUT_SECS |
admin.long_request_threshold_secs ⟳ | integer | 30 | seconds after which a running REST transfer appears in the console sessions; 0: all |
[[admin.roles]]
| Key | Type | Default | Effect |
|---|---|---|---|
admin.roles | array of tables | [] | the named admin roles |
admin.roles[].name | string | (required) | the name an authority carries (local account or JWT); unique, distinct from the [[roles]] names |
admin.roles[].permissions | list | (required, not empty) | among overview, sessions, kick, bans, unban, config, logs, audit, revoke |
[admin.session]
See Console sessions.
| Key | Type | Default | Effect |
|---|---|---|---|
admin.session | table | per-process key, 1 h, 12 h, revocations in memory | the admin console session tokens, signed by the server, and their revocations |
admin.session.key_file | path | absent | the signing key, 32 bytes at least (head -c 32 /dev/urandom), readable by the server only, the same on every instance; without it or secret_name, a key drawn for the process |
admin.session.secret_name | string | absent | the Kubernetes Secret whose admin-session.key entry holds the key; the chart creates it empty, the first pod writes it, the others read it (feature k8s); CRAFT_FILE_GATE_ADMIN_SESSION_SECRET_NAME; not with key_file |
admin.session.ttl_secs | integer | 3600 | life of a token; 60 to 86400 |
admin.session.max_age_secs | integer | 43200 | no renewal beyond this, counted from sign-in; 60 to 604800, at least ttl_secs |
admin.session.persist_file | path | absent | the file where revocations are shared, for them alone, distinct from the ban files; neither it nor backend: in memory, one instance, lost at restart |
admin.session.backend | string | file with persist_file, memory otherwise | file (with persist_file) or configmap (feature k8s), a ConfigMap like the bans; CRAFT_FILE_GATE_ADMIN_SESSION_BACKEND |
admin.session.revocation_configmap_name | string | craft-file-gate-access | the revocations ConfigMap, key revocations, with backend = "configmap"; the same as temporary accesses; CRAFT_FILE_GATE_ADMIN_SESSION_REVOCATION_CONFIGMAP_NAME |
admin.session.reread_interval_secs | integer | 5 | re-read of the shared revocations, in seconds; 1 to 30 |
[admin.grants]
Temporary accesses: a file role given to a user until a date, through POST /admin/grants (permission grant).
| Key | Type | Default | Effect |
|---|---|---|---|
admin.grants | table | in memory, 7 days at most | where temporary accesses are shared, and the longest one |
admin.grants.persist_file | path | absent | the file where accesses are shared, for them alone (no bans, no revocations); neither it nor backend: in memory, one instance, refused with [cluster] |
admin.grants.backend | string | file with persist_file, memory otherwise | file or configmap (feature k8s); CRAFT_FILE_GATE_ADMIN_GRANTS_BACKEND |
admin.grants.grant_configmap_name | string | craft-file-gate-access | the ConfigMap, key grants, next to the revocations; CRAFT_FILE_GATE_ADMIN_GRANTS_GRANT_CONFIGMAP_NAME |
admin.grants.reread_interval_secs | integer | 5 | re-read of the shared accesses, in seconds; 1 to 30 |
admin.grants.max_duration_secs | integer | 604800 | maximum duration of an access; 60 to 2592000 |
[admin.tls]
| Key | Type | Default | Effect |
|---|---|---|---|
admin.tls | table | absent: HTTP | HTTPS on the whole listener |
admin.tls.cert_file | path | absent | PEM certificate, re-read when the file changes; CRAFT_FILE_GATE_ADMIN_TLS_CERT |
admin.tls.key_file | path | absent | PEM private key; CRAFT_FILE_GATE_ADMIN_TLS_KEY |
admin.tls.auto_generate | boolean | false | generate a self-signed certificate (trials) |
admin.tls.auto_generate_cn | string | localhost | its Common Name |
admin.tls.auto_generate_dir | path | . | where to write it |
admin.tls.auto_generate_validity_days | integer | 31 | its validity in days |
[admin.ban]
Same keys as [sftp.ban]. See Bans and rate limits.
| Key | Type | Default | Effect |
|---|---|---|---|
admin.ban | table | absent: no ban | ban of addresses after authentication failures on the admin API; the file API has its own list, [api.ban] |
admin.ban.max_failures | integer | 5 | failures in the window before the ban; at least 1 |
admin.ban.ban_duration_secs | integer | 600 | duration of a ban in seconds, from the verdict |
admin.ban.window_secs | integer | 300 | failure counting window in seconds, fixed, opened by the first failure of a series |
admin.ban.whitelist_ips | list | [] | addresses and CIDR networks, IPv4 or IPv6, never banned by this instance |
admin.ban.trusted_proxies | list | [] | proxies whose X-Forwarded-For gives the client address; without control_listen, equal to api.ban.trusted_proxies |
admin.ban.persist_file | path | absent: in memory | file where bans are kept and shared between instances |
admin.ban.backend | string | "file" | "file", or "configmap" to share bans between Kubernetes pods |
admin.ban.ban_configmap_name | string | craft-file-gate-bans | the bans ConfigMap, with backend = "configmap" |
admin.ban.reread_interval_secs | integer | 5 | re-read of the shared persist_file in seconds, on top of watching; 1 to 30 |
[admin.metrics_thresholds]
See Metrics.
| Key | Type | Default | Effect |
|---|---|---|---|
admin.metrics_thresholds | table | - | alert thresholds of the console’s Metrics tab |
admin.metrics_thresholds.hash_pool_percent | integer | 80 | slots taken in the hashing pool, in % |
admin.metrics_thresholds.tls_cert_days | integer | 14 | days left on the admin TLS certificate, below |
admin.metrics_thresholds.rejections_per_minute | integer | 60 | SFTP refusals per minute |
admin.metrics_thresholds.jwt_refusals_per_minute | integer | 60 | JWTs refused per minute |
admin.metrics_thresholds.jwks_age_secs | integer | 2 x jwks_refresh_interval_secs | age of the JWKS cache |
admin.metrics_thresholds.clock_skew_secs | integer | 30 | clock skew of a storage |
admin.metrics_thresholds.cpu_percent | integer | 90 | CPU of the process, 100 = one core |
[api]
See Doors.
| Key | Type | Default | Effect |
|---|---|---|---|
api | table | absent: no API | the REST file API, on admin.listen |
api.enabled | boolean | false | serve the API, on the [admin] listener; requires [admin] |
api.prefix | string | /api/v1/files | path under which the API is served |
api.cors_origins | list | absent: no CORS | origins allowed for CORS; "*": all (see Bans) |
api.openapi | boolean | false | serve the Swagger UI (/api/docs) and the OpenAPI document (/api/openapi.json), without authentication; requires api.enabled |
[api.ban]
Same keys as [sftp.ban]. See Bans and rate limits.
| Key | Type | Default | Effect |
|---|---|---|---|
api.ban | table | absent: no ban | ban of addresses after authentication failures on the file API, its tickets and the file explorer |
api.ban.max_failures | integer | 5 | failures in the window before the ban; at least 1 |
api.ban.ban_duration_secs | integer | 600 | duration of a ban in seconds, from the verdict |
api.ban.window_secs | integer | 300 | failure counting window in seconds, fixed, opened by the first failure of a series |
api.ban.whitelist_ips | list | [] | addresses and CIDR networks, IPv4 or IPv6, never banned by this instance, nor capped in the hashing queue |
api.ban.trusted_proxies | list | [] | proxies whose X-Forwarded-For gives the client address (ban, limiter, audit); without admin.control_listen, equal to admin.ban.trusted_proxies |
api.ban.persist_file | path | absent: in memory | file where bans are kept and shared between instances |
api.ban.backend | string | "file" | "file", or "configmap" to share bans between Kubernetes pods |
api.ban.ban_configmap_name | string | craft-file-gate-bans | the bans ConfigMap, with backend = "configmap" |
api.ban.reread_interval_secs | integer | 5 | re-read of the shared persist_file in seconds, on top of watching; 1 to 30 |
[api.rate_limit]
See Bans and rate limits.
| Key | Type | Default | Effect |
|---|---|---|---|
api.rate_limit | table | absent: no limit | token bucket per address for the API |
api.rate_limit.requests_per_minute | integer | (required) | sustained request rate per address; at least 1 |
api.rate_limit.burst | integer | requests_per_minute | requests accepted in a row; at least 1 |
[api.ui]
See File explorer.
| Key | Type | Default | Effect |
|---|---|---|---|
api.ui | table | - | the web file explorer |
api.ui.enabled | boolean | false | serve the file explorer; requires [api] |
api.ui.path | string | /files | where the page is served, on the API listener (admin.listen, even with control_listen); starts with /, no trailing /; neither on api.prefix, nor on a server route (/admin, /ui, /metrics, /health, /livez, /readyz, /api/docs, /api/openapi.json) |
Reference: [log], [telemetry] and the rest
⟳: reloaded at runtime; no mark: taken at restart. See Hot reload.
[log]
| Key | Type | Default | Effect |
|---|---|---|---|
log | table | - | the logs and the audit trail |
log.level ⟳ | string | info | trace, debug, info, warn, error or off; CRAFT_FILE_GATE_LOG_LEVEL (or RUST_LOG), if set, wins, on reload too; see Log level |
log.format | string | json | json (one JSON line per event) or pretty (readable) |
log.audit ⟳ | string | all | volume of the trail: all, changes or failures; refusals and errors always written |
log.dir | path | absent: stdout only | directory of the log and audit files, rotated and compressed; CRAFT_FILE_GATE_LOG_DIR |
log.stdout | boolean | true | also write to stdout; false requires dir |
log.max_file_size_mb | integer | 100 | rotation before this size, in MiB; 1 to 1048576 |
log.retention_days | integer | 7 | days the archives of the application log are kept; 1 to 36500 |
log.audit_retention_days | integer | 90 | days the archives of the audit trail are kept; 1 to 36500 |
log.refusal_summary_threshold ⟳ | integer | 10 | identical refusals written per window before a summary line; 1 to 1000000 |
log.refusal_summary_window_secs ⟳ | integer | 60 | duration of that window; 1 to 86400 |
log.refusal_summary_max_addresses ⟳ | integer | 10000 | distinct refusals tracked at once; 1 to 1000000 |
[telemetry]
See Telemetry.
| Key | Type | Default | Effect |
|---|---|---|---|
telemetry | table | absent | the OpenTelemetry (OTLP) export |
telemetry.enabled | boolean | false | export traces |
telemetry.otlp_endpoint | URL | http://localhost:4317 | the OTLP collector, an http:// or https:// URL |
telemetry.service_name | string | craft-file-gate | the service.name attribute; service.version is the binary’s version |
telemetry.protocol | string | grpc | grpc (OTLP/gRPC, port 4317) or http (OTLP/HTTP protobuf, port 4318) |
telemetry.metrics | boolean | false | also export metrics over OTLP, to the same collector |
telemetry.metrics_interval_secs | integer | 60 | period of that export, in seconds; 1 to 3600 |
telemetry.on_exporter_error | string | refuse | an exporter that cannot be built: refuse (startup refused) or warn (startup without that signal) |
[uploads]
See Timeouts.
| Key | Type | Default | Effect |
|---|---|---|---|
uploads | table | - | upload timeouts, on all doors |
uploads.idle_timeout_secs | integer | 30 | upload abandoned after this time without a byte; 1 to 86400; CRAFT_FILE_GATE_UPLOAD_IDLE_TIMEOUT_SECS |
uploads.min_rate_bytes_per_sec | integer | 0: disabled | minimum average rate of an upload since its start, required once the grace has passed; at most 1073741824 |
uploads.min_rate_grace_secs | integer | 60 | wait before requiring that rate; 1 to 86400 |
uploads.takeover_idle_secs ⟳ | integer | 10 | an upload with no byte for this time is taken over by a new upload from the same account to the same file, on this instance; acts only below idle_timeout_secs (otherwise WARN); 0: never; 0 to 86400; on reload, an out-of-bounds value keeps the one in force |
[uploads.stale_partials]
See Atomic writes.
| Key | Type | Default | Effect |
|---|---|---|---|
uploads.stale_partials | table | - | sweep of the leftovers of interrupted uploads |
uploads.stale_partials.age_check | boolean | true | measure age on the storage clock (a probe file in the same directory), not on the server’s |
uploads.stale_partials.grace_secs | integer | 2 x idle_timeout_secs + max(300, idle_timeout_secs), i.e. 360 s | age from which a leftover is deleted; more than 2 x idle_timeout_secs + 60, at most 2592000 (30 days) |
[tcp_keepalive]
| Key | Type | Default | Effect |
|---|---|---|---|
tcp_keepalive | table | - | TCP keepalive of every accepted connection |
tcp_keepalive.enabled | boolean | true | enable SO_KEEPALIVE |
tcp_keepalive.idle_secs | integer | 30 | silence before the first probe; 1 to 32767 |
tcp_keepalive.interval_secs | integer | 10 | gap between two probes; 1 to 32767 |
tcp_keepalive.count | integer | 3 | unanswered probes before closing; 1 to 127 |
tcp_keepalive.user_timeout_secs | integer | 0: derived from idle_timeout_secs | Linux: how long sent bytes can stay unacknowledged, then the connection is closed; 5 to 86400 |
[reload]
| Key | Type | Default | Effect |
|---|---|---|---|
reload | table | - | how a modified file is noticed |
reload.watch | string | auto | auto: inotify, and periodic re-reading if inotify cannot be used; inotify: inotify required; poll: periodic re-reading only, no inotify instance |
reload.poll_interval_secs | integer | 5 | period of the periodic re-reading, in seconds; 1 to 60 |
[security]
See Security.
| Key | Type | Default | Effect |
|---|---|---|---|
security | table | - | how trust files are judged |
security.allow_group_writable_trust_anchors | boolean | false | a trust file writable by the server’s group: WARN instead of a refusal; never the configuration file |
Reference: environment variables
The CRAFT_FILE_GATE_* variables replace a key from the file, after the file
is read. They are mostly useful in containers.
Variables that replace a key
| Variable | Key replaced | Value |
|---|---|---|
CRAFT_FILE_GATE_LISTEN | sftp.listen | ip:port; with [sftp] |
CRAFT_FILE_GATE_PROBES_LISTEN | server.probes_listen | ip:port |
CRAFT_FILE_GATE_ADMIN_LISTEN | admin.listen | ip:port |
CRAFT_FILE_GATE_ADMIN_CONTROL_LISTEN | admin.control_listen | ip:port |
CRAFT_FILE_GATE_ADMIN_HEADER_READ_TIMEOUT_SECS | admin.header_read_timeout_secs | integer, 1 to 300 |
CRAFT_FILE_GATE_ADMIN_BEARER_TOKEN | admin.bearer_token | the token |
CRAFT_FILE_GATE_ADMIN_BEARER_TOKEN_FILE | admin.bearer_token | a file that holds the token |
CRAFT_FILE_GATE_ADMIN_TLS_CERT | admin.tls.cert_file | path; creates [admin.tls] if missing |
CRAFT_FILE_GATE_ADMIN_TLS_KEY | admin.tls.key_file | path; creates [admin.tls] if missing |
CRAFT_FILE_GATE_ADMIN_SESSION_SECRET_NAME | admin.session.secret_name | the Secret name; the chart sets it |
CRAFT_FILE_GATE_ADMIN_SESSION_BACKEND | admin.session.backend | file or configmap; the chart sets it next to the shared Secret |
CRAFT_FILE_GATE_ADMIN_SESSION_REVOCATION_CONFIGMAP_NAME | admin.session.revocation_configmap_name | the name of the revocations ConfigMap; the chart sets it |
CRAFT_FILE_GATE_ADMIN_GRANTS_BACKEND | admin.grants.backend | file or configmap; the chart sets it |
CRAFT_FILE_GATE_ADMIN_GRANTS_GRANT_CONFIGMAP_NAME | admin.grants.grant_configmap_name | the name of the temporary access ConfigMap; the chart sets it |
CRAFT_FILE_GATE_SFTP_BAN_CONFIGMAP_NAME | sftp.ban.ban_configmap_name | the name of the bans ConfigMap; the chart sets it |
CRAFT_FILE_GATE_API_BAN_CONFIGMAP_NAME | api.ban.ban_configmap_name | likewise |
CRAFT_FILE_GATE_ADMIN_BAN_CONFIGMAP_NAME | admin.ban.ban_configmap_name | likewise |
CRAFT_FILE_GATE_CLUSTER_LISTEN | cluster.listen | ip:port; creates [cluster] if missing; the chart sets it |
CRAFT_FILE_GATE_CLUSTER_PEERS | cluster.peers | dns:<name>:<port>, or comma-separated <host>:<port> entries; the chart sets it |
CRAFT_FILE_GATE_CLUSTER_SECRET_NAME | cluster.secret_name | the Secret name; the chart sets it |
CRAFT_FILE_GATE_CLUSTER_MIN_PEERS | cluster.min_peers | integer; the chart sets it (cluster.minPeers) |
CRAFT_FILE_GATE_CLUSTER_UNREADY_WHEN | cluster.unready_when | comma-separated detectors; the chart sets it with cluster.unreadyWhen |
CRAFT_FILE_GATE_JWT_SECRET | auth.jwt.secret | the HMAC secret |
CRAFT_FILE_GATE_JWT_SECRET_FILE | auth.jwt.secret | a file that holds the secret |
CRAFT_FILE_GATE_HASH_WORKERS | auth.hash_workers | integer, 1 to 1024 |
CRAFT_FILE_GATE_HASH_QUEUE | auth.hash_queue | integer, 1 to 65536 |
CRAFT_FILE_GATE_LOG_LEVEL | log.level | trace … off |
CRAFT_FILE_GATE_LOG_DIR | log.dir | path |
CRAFT_FILE_GATE_UPLOAD_IDLE_TIMEOUT_SECS | uploads.idle_timeout_secs | integer, 1 to 86400 |
A variable applies to the key it replaces, where that key has an effect
(CRAFT_FILE_GATE_ADMIN_* with [admin], CRAFT_FILE_GATE_JWT_SECRET with
an HMAC algorithm). Once CRAFT_FILE_GATE_LOG_LEVEL is set, the file no
longer sets the level. The startup line config override from env names each
variable applied, never showing a secret.
Secrets: the _FILE form
CRAFT_FILE_GATE_ADMIN_BEARER_TOKEN_FILE and CRAFT_FILE_GATE_JWT_SECRET_FILE
name a file read once, at startup. Prefer them: a variable
stays readable in /proc/<pid>/environ.
| Rule | Effect |
|---|---|
| the file | a trust file, in UTF-8, not empty; only one of the two forms X and X_FILE |
| a trailing newline | removed |
| file modified | taken at the next restart |
A Kubernetes Secret mounted read-only works as is: see Kubernetes secrets.
Other variables read
| Variable | Effect |
|---|---|
RUST_LOG | replaces the log filter; keep audit=info in it, otherwise the audit trail goes silent (WARN at startup) |
OTEL_EXPORTER_OTLP_PROTOCOL, OTEL_EXPORTER_OTLP_TRACES_PROTOCOL | read to report in a WARN that they contradict [telemetry] protocol; the file decides |
SSL_CERT_FILE, SSL_CERT_DIR | the TLS roots, if the image carries none (:scratch) |
TOKIO_WORKER_THREADS | threads of the main runtime; quoted on the startup line |
HOSTNAME | the pod_name of sessions in the admin API |
Reference: audit line vocabulary
Each audit line is an event of the audit target, one JSON line per event.
Reading and filtering the trail: Audit. What an error
means and where to fix it: Troubleshooting.
Actions
action | source | Written when |
|---|---|---|
list | sftp, api | listing of a directory |
download | sftp, api | end of the transfer |
upload | sftp, api | end of the transfer |
mkdir | sftp, api | one line per directory created, from outermost to innermost |
delete | sftp, api | SFTP REMOVE of a file; REST DELETE of a file or an empty directory |
rmdir | sftp | RMDIR of an empty directory |
delete_recursive | api | DELETE of a non-empty directory |
rmdir_recursive | sftp | RMDIR of a non-empty directory |
rename | sftp, api | rename or move |
stat | sftp, api | refusals only; REST HEAD |
setstat | sftp | refusals only |
request | api | request refused before the handler with no route for its method, or GET ?rights |
connection_accepted | sftp, api, admin | successful authentication: one per SFTP session, one per REST request, one per 15 min admin window |
connection_rejected | sftp, api, admin | refused authentication |
connection_rejected_summary | sftp, api, admin | repeated anonymous refusals, summarized at the end of the window |
ip_banned | sftp, api, admin | an address is banned, once per switch |
session_end | sftp | an SFTP session leaves the registry |
session_kick | admin | DELETE /admin/sessions/{id} |
unban | admin | DELETE /admin/bans/{protocol}/{ip} |
logout | admin | POST /admin/logout under a JWT or the static token: nothing is revoked |
session_revoke | admin | POST /admin/revocations, POST /admin/logout under a session token: username the revoked account, admin the author |
grant_add, grant_revoke | admin | POST /admin/grants, DELETE /admin/grants/{id}: username the beneficiary, role, grant_id, until, grant_reason (the reason given), lift |
audit_read | admin | reading of the trail by the console or its stream |
get_status, get_config, get_resources, list_sessions, session_roles, logs_status, get_logs, events, stream_logs, list_bans, list_grants, grant_candidates | admin | insufficient permission refusal of a read route |
Results
result | Meaning |
|---|---|
success | the operation succeeded |
denied | the server refused it |
error | the storage made it fail, or nothing could be built for it, or it was interrupted |
unknown | the server does not know what became of it (interrupted commit) |
Fields
| Field | Lines | Meaning |
|---|---|---|
source | all | the door: sftp, api, admin; for ip_banned, the ban list |
action, result | all | above |
reason | denied, error, unknown; absent on success | a fixed string, below |
username | all except summaries | the name presented; empty for an unverified token |
remote_addr | all | the client IP; with the port on SFTP connection_accepted and session_end |
session_id | SFTP operations, session_end, session_kick | the session identifier |
path | operations | the path as the user sees it; the source of a rename |
new_path | rename | the destination |
backend | operations | the name of the mount’s backend; "" on a synthetic path |
count | upload, download, list | bytes transferred, or entries listed; 0 elsewhere |
replaced | upload | yes, no, unknown: did the upload destroy existing content |
removed | delete_recursive, rmdir_recursive | entries removed; unknown on a local backend |
took_over | upload | the stalled upload of the same account that this one took over: its SFTP session_id, or its REST transfer_id; "" elsewhere |
transfer_id | REST upload | the transfer identifier, the one the console shows; "" elsewhere |
grant_id | operations | the temporary access that brought the mount of the path (several: comma-separated); "" elsewhere |
auth_method | connection_* | SFTP password, pubkey, jwt; API basic, bearer, ticket; admin jwt, static_token, session, password (console sign-in); none without credential |
signature_algorithm | SFTP connection_accepted | algorithm of the key’s signature |
suppressed, threshold, window_secs, overflow | connection_rejected_summary | refusals not written, and the window settings |
ban_duration_secs, expires_at_epoch | ip_banned | duration and end of the ban |
duration_secs, bytes_read, bytes_written | session_end | duration and volume of the session |
admin, admin_addr, admin_auth_method | admin actions | the author |
kicked_by | session_kick | the author of the kick |
protocol, lift | unban | the list (sftp, api, admin); held, local_only, overruled, not_banned |
lift | session_revoke | held, local_only |
permission | insufficient permission refusals | the missing permission |
lines, levels, since, until, q, fields | audit_read | the read performed; lines = 0 for a stream |
A field that does not apply is empty ("", 0), not absent. No line carries
a password, a token or the text of a storage error: that text goes to the
application log, alongside.
reason values
File operations:
reason | result | Meaning |
|---|---|---|
acl | denied | the ACL does not grant the right |
acl subtree | denied | deletion of a tree that the ACL does not fully cover |
synthetic path | denied | write on a synthetic directory or a mount point |
rename across mounts | denied | source and destination under two mounts |
invalid path | denied | control character, \, forbidden Windows form |
reserved name | denied | name the server reserves for itself (upload lock) |
rename into restricted | denied | rename that would drop content without write |
range not satisfiable | denied | REST download whose Range starts after the end of the file (416) |
symlink escape | denied | symbolic link outside the local root |
exists | denied | existing destination, or exclusive creation on a taken name |
is a directory, not a directory | denied | REMOVE of a directory, RMDIR of a file |
upload in progress | denied, error | another upload holds the destination |
taken over | error | stalled upload taken over by a new upload of the same account (uploads.takeover_idle_secs) |
quota exceeded (and , truncated file removed, , append not undone) | denied | max_file_mb exceeded |
session killed | denied | session cut during the operation |
no roles, username not usable as home directory | denied | REST: no role, unusable name |
role resolution, mount conflict | error | REST: roles not resolved, conflicting mounts |
not found, permission denied, already exists, not a directory, is a directory, directory not empty, storage error, not implemented, unsupported, upload in progress | error | kind of the storage error |
session ended: <cause> | error | SFTP session ended during the operation; <cause> is the reason of the session_end |
session ended | error | REST client left during the request |
upload idle timeout, upload below minimum rate | error | upload cut by the server |
commit interrupted | unknown | client left while an upload was being published |
Authentication (connection_rejected, connection_rejected_summary):
reason | result | Meaning |
|---|---|---|
| absent | denied | wrong password, unknown name, SFTP key or JWT refused: nothing tells which accounts exist |
no authorized key offered | denied | no offered key was authorized |
signature algorithm not allowed | denied | signature algorithm of the key not allowed |
missing credential, empty credential, malformed credential | denied | Authorization header absent, empty, unreadable |
invalid token, expired token, wrong issuer, wrong audience | denied | token refused by the verifier |
verifier unavailable | denied, error | no key matching the token, or nothing to verify it |
no username claim | denied | verified token without a name |
method disabled | denied | authentication method turned off |
no matching roles, no matching admin roles | denied | credential accepted, no role |
role resolution, mount conflict, backend initialization failed | error | credential accepted, session impossible to build |
username not usable as home directory | denied | the name cannot be a directory |
session limit | denied | max_sessions_per_user reached |
banned, rate limit | denied | address banned, rate exceeded |
password checks saturated, password checks saturated for address, shutting down | error | password not verified: pool full, address share reached, shutdown |
invalid ticket, expired ticket, revoked ticket | denied | download ticket refused |
revoked token | denied | admin console session token whose account was removed or changed its password, or was revoked |
Session end (session_end): closed, connection_ended, admin_kick,
shutdown_idle, shutdown, banned, internal_error, grant_expired,
grant_revoked (a temporary access that the connection used has ended:
grant_id names it; a REST transfer cut this way says
session ended: grant_expired or grant_revoked).
Admin actions: insufficient permission, session not found, no sessions,
not banned, overruled, unknown protocol, invalid IP address,
no ban manager.
Migrating from ProFTPD
Moving from a ProFTPD with virtual users (AuthUserFile, ftpd.passwd)
to CraftFileGate.
The model
ProFTPD puts identity, home and rights in one ftpd.passwd entry and in
<Directory> blocks. CraftFileGate separates them: users.toml says who you
are (username, password_hash, authorized_keys, authorities),
roles.toml what you see ([[backends]], [[roles]] and their
mounts). The home belongs to a role’s mount, not to the user:
Users, roles and mounts.
Translating a ftpd.passwd entry
alice:$6$rounds=5000$xyz...:1001:1001:Alice:/srv/ftp/alice:/bin/false
# users.toml
[[users]]
username = "alice"
password_hash = "$6$rounds=5000$xyz..." # kept as is, see below
authorized_keys = ["ssh-ed25519 AAAA... alice@poste"]
authorities = ["clients"]
# roles.toml: one role for all clients
[[backends]]
name = "ftp"
type = "local"
root = "/srv/ftp"
[[roles]]
name = "clients"
[[roles.mounts]]
backend = "ftp"
home_dir = "/{username}" # DefaultRoot ~: alice sees /srv/ftp/alice as /
create_home = true # CreateHome on
acl = [{ path = "/", rights = ["read", "write", "list", "delete", "rename"], recursive = true }]
| ProFTPD field | Becomes |
|---|---|
| password | password_hash |
| uid, gid | nothing: everything is written by the service account, without setuid |
| home | the mount’s home_dir, {username} for %u |
| shell, gecos | nothing: only the sftp subsystem is served |
{username} accepts only A-Z a-z 0-9 . _ -, 64 characters at most, with no
leading dot; any other name cannot connect to that mount.
Passwords
[auth.methods]
local = { enabled = true, allow_sha512_crypt = true } # or allow_bcrypt
$6$ and bcrypt are taken as is (Hashes);
DES and MD5 ($1$) need a new password or a public key.
Keeping the hashes is a migration step: sha512-crypt is not
memory-hard. Regenerate them as argon2id at the first opportunity:
echo -n "mot-de-passe" | craft-file-gate hash-password --user alice # a [[users]] block
An account that logs in only by key keeps a password_hash (the field is
required): put the hash of a random secret in it, or turn off passwords
for everyone (local.password = false). An identity provider can
also replace users.toml: see Authentication.
Directive mapping
| ProFTPD | CraftFileGate |
|---|---|
AuthUserFile | auth.methods.local.users_file |
AuthGroupFile | authorities and [[roles]] |
DefaultRoot ~ | home_dir = "/{username}" |
CreateHome on | create_home = true (local backend) |
<Directory> and <Limit> | [[roles.mounts.acl]], see below |
MaxLoginAttempts | [sftp.ban] max_failures: bans the address |
MaxClientsPerUser | server.max_sessions_per_user |
MaxClients | nothing; [sftp.rate_limit] bounds the connection rate |
HiddenStores on | [server.hidden_stores] enabled = true; disabled by default, as in ProFTPD |
mod_quotatab | the mount’s max_file_mb: a cap per file, not a cumulative volume |
TransferLog, ExtendedLog | the audit trail, see Audit |
TLSEngine (FTPS), PassivePorts, MasqueradeAddress | nothing: SFTP on a single TCP port |
SFTPHostKey | sftp.host_keys |
SFTPAuthorizedUserKeys | the user’s authorized_keys |
mod_sql, mod_ldap | [auth.jwt] or auth.authz_base_url |
Umask | the umask inherited at launch, for client files (Security) |
<Anonymous> | a dedicated account, with a read-only role |
With hidden_stores, the in-progress file is named as in ProFTPD, with
a token: .in.rapport.csv.<token>. (Atomic writes).
Translating <Limit> blocks
| FTP commands | Right |
|---|---|
RETR, READ | read |
STOR, APPE, MKD, WRITE | write |
LIST, NLST, CWD, DIRS | list |
DELE, RMD | delete |
RNFR, RNTO | rename |
There is no inheritance: the most specific entry decides alone, and a
read-only subtree (<Limit WRITE> DenyAll) restates everything still
allowed there, { path = "/archive", rights = ["read", "list"], recursive = true }.
A path that no entry covers is refused. See ACL.
The cutover
| Topic | CraftFileGate |
|---|---|
upload resume (AllowStoreRestart) | served, except with hidden_stores and on S3 |
| file ownership | the service account; any other ownership goes through the directory’s permissions (setgid, POSIX ACL) |
| FTP and FTPS | not served: each client moves to SFTP or to the REST API |
- Inventory the clients; extract names, hashes and homes from
ftpd.passwd. - Write
roles.tomlandusers.toml; launch on another port. - Validate with a test account (each right, one refusal), distribute the secrets.
- Shut down ProFTPD, switch the port, follow the audit trail.