The self-hostable companion to HelloID. A dependency-free web app that works alongside the platform: it analyses what is there — risk, cost, coverage, unmanaged access — and prepares what comes next — mined and condensed business rules, account matching, naming conventions, field-mapping simulations. Drop any HelloID export on the page and it shows the most analysis that export can carry — every import is optional, and each one deepens the model rather than gating it.
The reconciliation (ReconciliationReport.csv) is the base of the access analysis: the
account ↔ entitlement ↔ person graph, risk scoring, licence drift pricing, and the diff
against any earlier import. The vault (people and contracts) powers the people and
organisation views on its own — org walker, department scorecards, workforce analytics,
attestation packs — and, joined with the reconciliation, makes rule conditions evaluable
and separates leavers from unowned accounts. The business rules split entitlements into
modelled, draft-only and unmodelled; granted entitlements and historic actions
distinguish "granted outside HelloID" from "granted by HelloID and not recorded", and turn
a missing entitlement into a failed grant on a date; the product catalogue and
product assignments connect Service Automation requests and approvals to the access
they explain — helloid-export.py / helloid-export.ps1 pull both from the tenant into one
helloid-service-automation.json (either half may be empty); the two older files still load.
A view whose exports are missing is not hidden: it wears a lock in the sidebar and renders a page naming exactly which export unlocks it, what that export contains and where in HelloID it lives. Load only a vault and the people views work; add the reconciliation later and everything merges — the model is rebuilt from whatever is loaded, so import order never changes the outcome.
Serve it from a laptop or any static host — it is plain files, with no build step, no CDN and no back end. Wherever it runs, the exports are read and analysed in the browser that opened the page: they are never uploaded to the host serving it.
Sidekick does three jobs for three audiences, and the first thing a session asks is which
one it is for (#choose, or ?edition=govern|health|consult; the hostname prefix
helloid-govern. etc. presets it):
- Govern — governance analytics for partners and their customers: compliance against NIS2 / ISO 27001 / BIO 2.0, risk, cost, evidence of decisions, one person end to end, the board pack.
- Health check — is this HelloID well built and does it run: rules against reality, what the engine did and failed, naming conventions, correlation quality, what to condense, add or retire.
- Consult — design the model before HelloID exists: roles from HR, naming conventions, field-mapping simulation, the Nedap ONS workbench, matching.
An edition is a sidebar and a landing page, not a build: every view stays reachable by its
hash and says which edition it belongs to; ?edition=all shows everything.
A partner runs Sidekick for several customers in one browser. The topbar picker holds one workspace per customer: snapshots, imports, settings, branding and the chosen edition are kept apart, so the board pack carries that customer's name and logo and the next customer does not overwrite the last. Switching reloads the page. The first workspace keeps the storage the app always used, so nothing migrates; the demo loads into its own workspace and never touches a customer's.
A dashboard is what one audience gets and nothing more, on a URL of its own: the
built-in HR dashboard lives at /d/hr (nginx rewrites the app's relative assets
back to the root; ?dashboard=hr does the same where no such rule exists and is
remembered — ?dashboard=off leaves it; an hr. / helloid-hr. hostname presets it too).
It shows People with Person 360 and Organisation — no imports, no settings, no workspace picker, no money, no risk scores.
It cannot authenticate anyone: it is a preset, and who may open the hostname is the
proxy's business (Cloudflare Access, for instance).
HR imports nothing. The dashboard reads a published bundle: on Data points › Publish
for a dashboard the partner exports one file with every data point (vaults included) and
the rules and activity loaded now, puts it in published/ on the host and names it in
published/dashboards.json (see docs/dashboards.example.json). The dashboard fetches it
on open and takes a newer bundle (higher exportedAt) over what the browser holds; "Reload
data" fetches again. published/ is mounted read-only into the container and is never in
git or the image. A dashboard runs in a workspace of its own, so the partner's data in the
same browser is untouched. Later the source may be a HelloID API that answers the same
shape; the loader will not need to change. dashboards.json can also define other
dashboards: views, landing, role, source (and a legacy hide list that narrows
the role further).
Roles and facets. What an audience may see is decided in one place (js/access.js),
not view by view. Every sensitive thing belongs to a facet — money (prices, spend,
€/head, cost tabs), risk (scores, bands, bars, outliers, why-score), governance
(findings, controls, SoD, attestation, exclusions, evidence), audit (who did what,
engine health), admin (imports, settings, data points, collectors) — and a role says
which facets it has: full (everything), hr (none: identity, contracts, accounts and
access facts only), manager (governance), auditor (risk, governance, audit). The
shared building blocks enforce it: tiles, table columns, tabs and cards carry a facet and
are not drawn without it; the account and permission drawers show only the facts; the
money and score formatters print a dash for a role without the facet, whoever calls them;
and after every render and drawer a sweep removes anything marked data-facet for a
hidden facet. A browser leak test walks every view, tab and drawer of the HR dashboard
and fails on any €, score bar or governance label. The app still authenticates nobody: a
role is what a dashboard shows, and who may open the dashboard is the host's business.
make-demo-set.py writes a fictional organisation into demo/: nine exports describing
the same invented people — reconciliation, vault (employer › divisions › departments),
a directory collection, rules, granted and historic actions, the Service Automation
catalogue and assignments, and a HelloID audit log — so every file joins every other
and the fit check reads them as one tenant. Where that directory is published, the
Imports view offers to load them all in one click.
While it is loaded, a banner sits above every view and every page of the board report
carries a printed mark, because a PDF outlives the tab it came from.
python3 make-demo-set.py # -> demo/*.csv, demo/vault.json, demo/manifest.json
python3 make-demo-set.py --rows 20000 # a larger tenantEverything in it is invented — names from a fixed word list, employee numbers counting up from 500000, avo.local as the domain. It is the only CSV/JSON the shipped nginx config is allowed to serve; every other export path is denied.
./serve.sh # http://localhost:8123Opening index.html (or the bundled single file) directly from disk also works, but
browsers give file:// pages an opaque origin, so snapshots cannot be persisted and the
sample loader is hidden. Over http both work.
Import in the top bar opens a slot per export — what it is, what it unlocks, where in HelloID it comes from, and whether it is loaded. A slot only accepts its own kind of file, and every slot can be replaced or removed independently; removing the reconciliation keeps the snapshot archive, so the Snapshots view can bring any earlier import back. Dragging a file anywhere onto the page still routes it on content.
No reconciliation export is committed to this repository — they carry account and person
names, so *.csv is gitignored. Drop your own export in this folder and the "load the export
from this folder" button picks it up (it looks for ReconciliationReport.csv, then
sample-recon.csv); drag-and-drop works regardless of where the file sits.
python3 make-sample.py # 5,000 rows -> sample-recon.csv
python3 make-sample.py --rows 50000 # for performance work
python3 make-sample.py --systems 2 # multi-system export
python3 make-sample.py --seed 7 -o next.csv # a second run to diff against the firstFictional names, HelloID-shaped group naming, and enough of every case that all fifteen finding rules have something to fire on: unowned admin accounts, disabled accounts still in licence groups, stacked SKUs, a security baseline at ~88% coverage, break-glass groups with two members, rows already dispositioned in HelloID. Output is deterministic per seed, so two seeds give you a pair of exports to try the diff on. The generated file is gitignored like any other export.
python3 build.pyproduces two things in dist/:
| Artefact | What it is | Use it when |
|---|---|---|
reconciliation-analytics.html |
~270 KB single file, every script, stylesheet and the logo inlined | mailing it, dropping it on a share, handing it to a customer. Double-click opens it; drag the CSV on. Snapshots stay in the tab unless it is served over http. |
helloid-sidekick-<version>.zip |
the folder form, plus serve.sh |
the customer wants diffing across sessions, or you want to keep editing it |
Both are static: no install, no runtime, no network access. If you want it permanently available for a team, the folder can also be dropped on any static host (an internal IIS vdir, S3, a share served over http) — there is no back end to deploy.
Versions are CalVer: YYYY.M.N — the Nth release of that month. The source of
truth is js/changelog.js (newest entry on top); the topbar version, the About
card in Settings, the in-app "What's changed" drawer, build.py's zip name and
CHANGELOG.md all read from it. When a change set ships: add an
entry at the top of js/changelog.js, then regenerate the markdown with
node make-changelog.jsEach release is also tagged and published on GitHub:
git tag -a v<version> -m "<version>"
git push origin --tags
gh release create v<version> --title "<version>" --notes "<the entry's bullets>" --latestEverything — interface, findings, board report and the Markdown export — is available in
English and Dutch, switchable from the top bar. Numbers, currency and dates follow the
selected language. The default follows the browser. All strings live in js/i18n.js;
adding a third language means adding one more block there.
The first view is a print-ready board report: A4 pages, plain language, no scores without a sentence explaining what they mean. Cover with the overall assessment, a one-page summary, a traffic-light table of themes, the money page, change since the previous review, a numbered recommendation list with effort and yearly saving per action, and a short method page. Export PDF opens the browser print dialog — choose Save as PDF, A4 portrait. Organisation, author and date are filled in above the sheets and are not printed.
Three slots, because a square app mark and a wordmark are not interchangeable:
| Slot | File | Used for |
|---|---|---|
| icon | assets/icon.svg | .png |
top bar, browser tab |
| logo | assets/logo.svg | .png |
wordmark on light backgrounds |
| logo (dark) | assets/logo-light.svg | .png |
wordmark on the report cover's dark chip |
Files dropped in assets/ are detected automatically and inlined into the single-file
bundle by build.py. Anything uploaded under Settings → Report branding is stored in
the browser as a data URI, which is what makes it survive into the printed PDF. Each slot
falls back to the next; with none at all the app shows a neutral gradient tile. The same
panel sets the product name shown in the title bar, the accent color (buttons, links,
active tabs, the first chart color, the report — both themes), the default theme a
fresh visitor gets (their own toggle wins afterwards) and a welcome line for the empty
start page. The Board report page adds a contact row on the cover and a footer
line printed on every sheet. All of it rides the settings export.
Dockerfile + docker-compose.yml serve the folder from nginx on loopback, fronted by a
Cloudflare tunnel — see DEPLOY.md. Only index.html, css/, js/ and assets/
enter the image; .dockerignore keeps exports out and nginx returns 403 for .csv/.json
outside /assets/ in case one is ever copied onto the host.
Nothing needs protecting server-side: the page ships no data, makes no network calls, and every visitor starts empty and drags in their own exports.
Everything stays client-side, and all of it is listed — with live sizes, per-store clear buttons and a clear-everything — under Settings → Storage:
| Store | Holds |
|---|---|
localStorage hr.config.v1 |
settings and decisions: taxonomy, prices, weights — and the match book, Nedap book, classification overrides and permission notes (these contain names from your data) |
localStorage hr.brand |
uploaded logos and the product name |
localStorage hr.nav.v1 / hr.theme / hr.lang |
UI preferences |
IndexedDB helloid-recon/snapshots |
one snapshot per reconciliation import, for the Diff view |
IndexedDB helloid-recon/context |
the raw text of every companion import (vault, directory, rules, activity, …) so a session survives a reload |
The same tab has the kill switch: Remember data in this browser off wipes all of the above and turns every save path into a no-op — the session keeps working in memory and forgets on reload. Only that preference itself persists, so the choice sticks. The first persisting import announces itself once with a pointer to this tab. Everything configurable also travels by file (settings, match book, Nedap book, snapshot bundle); note the settings file contains the decisions and books, names and employee ids included. Anonymous usage statistics (event names and size buckets, never data) have their own toggle on the same tab.
Every import stops at a review step before any number is shown. It reports how much of this export the current settings actually describe — permissions categorised, accounts classified, groups priced — and proposes rules mined from the export's own naming:
- permission categories from prefixes that no current rule matches, with a sensitivity
pre-filled from the prefix itself where it is recognisable (
BEH-,SRV-,SEC-…) - account classes from recurring name shapes. Marker vocabulary (
adm-,svc-,ext_) is proposed from two accounts up; anything else has to be common enough in the export to look deliberate.firstname.lastnameis treated as a naming convention, not a marker, so surnames are not proposed as classes. - groups that look priceable but have no price yet, so licence spend is not silently zero
Each proposal shows how many names it matches right now, is editable before it is applied, and nothing is written until you select it. Continue without changes skips the lot; a checkbox turns the step off for future imports (re-enable in Settings by resetting).
Classification workbench (#classify, where an import with unclassified names lands):
the rules are the engine and the fallout is the work list. Two tabs, entitlements and
accounts, each with the recognition rules live (hits per rule, and wins — the names it
is the first to claim; order decides, so a specific rule moves above a general one), a
fallout strip (unclassified / by a rule / overruled by hand / family answer; for accounts
also unknown name shape versus user by default — a plain firstname.lastname is a user
account), and a table of every name with the draft's answer, who decided it, and a
dropdown that overrules it. Select rows to set a category for all of them or to make a
rule from them (the word most of them share, or their common prefix). Old wizard
answers keep working, shown as family answer, convertible to a rule or forgotten.
Nothing counts until Save & re-score.
The "recognised" answers come from the recognition vocabulary on Settings › Recognition
(its own tab; Classification keeps the definitions — categories with their sensitivity,
account types with their weight, toxic pairs, workforce categories). Each entitlement row
says in plain language how its words match a name — starts with, contains, ends
with, is a word (split on separators and case changes, so adm hits GG_ADM_SQL but
not admissions), is exactly — and which category that means; the first row that hits
wins, a wizard answer beats every row, and every row shows how many names in the loaded
import it hits. The built-in table (vocabulary v2) matches whole words — split on
separators and case changes, digits after the word allowed — so log catches
LOG-Eventcollector and srv catches SRV01-Backup, while LOGISTIEK, BEHANDELAAR,
PRIVACY, SECRETARIAAT, NASCHOLING and SharePoint (for file shares) no longer
count. It opens with the directory's own privileged groups by exact name (Domain /
Enterprise / Schema Admins, Administrators, the Operators groups, the Entra administrator
roles), then privileged, server, security, business role, SharePoint/OneDrive, file share,
application (incl. Citrix, RDS/VDI, SAP, AFAS, D365, ECD), mailbox, project, distribution
list, team, device (incl. Intune, WLAN) and licence (incl. Entra SKUs such as SPE_E5,
ENTERPRISEPACK). Account-type rows match an account's cohort word exactly; a, sa and
gen are gone (they made a.jansen an admin). A workspace that stored the older table is
upgraded row by row — untouched built-in rows become the new ones, edited and added rows
stay — and the first build after the upgrade lists every name that moved, with what it was,
until Keep the update is pressed. New rows default to is a word.
One resolver. Every category and account type is decided in js/classify.js — the
model build, the workbench preview and the findings all call it, so what the workbench
shows is what Save produces. Order: item answer › family answer › rules › signals ›
fallback (other / user, by id). Rules apply to every name, prefixed or not; a rule
whose category was deleted is skipped rather than swallowing names. Item answers for
entitlements are kept per system. Account names lose their directory decoration first
(DOMAIN\, @domain, a trailing $); a recognised trailing word beats an unrecognised
leading one (jan-adm is an admin account); #EXT# guests are external and …$
managed service accounts are service accounts when no rule says otherwise. The
workbench saves only the parts you edited, and asks before replacing a part that was
changed elsewhere meanwhile.
Signals beyond the name. With a directory export loaded (AD or Entra), the group's own
facts count. Hard facts outrank a name rule: a licence SKU from Entra is a licence, a
group the directory calls Distribution is a distribution list (an AGDLP DL_… security
group is not), and a group nested at any depth in a privileged group is privileged — it
hands out what its parent has. Soft hints only place what no rule did: a Microsoft 365
(Unified) group is a team, and a price-book row written for licences that names a group
by pattern makes it a licence. Accounts in a service-account OU are service accounts, in
an admin/beheer OU admin accounts. A rule can also read where a group lives (lives in
(OU / path)): the DN in the reconciliation's brackets or the directory OU. In the
workbench, an unplaced entitlement held only by admin and service accounts is marked
probably privileged with a one-click confirm — a suggestion, never a silent decision.
How much is really placed. The Overview's Classified ring weighs by access: of every
assignment (an account holding an entitlement), the share whose entitlement has a category
and whose account a known type — an unplaced group with forty holders weighs forty, not one
name in three thousand. The count by name stays under it. Compliance warns below 80 %. A
privileged or server answer that rests on a short word alone (ADM-Finance by adm) is
marked weak and gets its own review filter in the workbench; the unclassified backlog
opens sorted by holders.
Workbench feedback. Make a rule picks the word that best separates the selection from
everything else (not the most common one — in a GG_* scheme every name shares gg), says
how many selected names it catches and how many others, and inserts the rule just above the
first rule that claims one of them. Save first lists every entitlement and account
whose answer changes; Undo last save puts back what the save replaced. Export
rules writes the recognition rows and family answers — not this tenant's per-name
answers — to a file another workspace can import into its draft.
Tests: NODE_OPTIONS= node test/classify.test.js runs the golden list against the real
modules (no dependencies).
A pattern tester sits on the same page and in Settings: type a regex, pick permission or account names, and see the match count, a sample of hits, and a warning when the pattern is invalid or matches everything. Every rule row in Settings also shows its live match count against the loaded import.
Every account gets a 0–100 score from named parts — an identity nobody owns, privileged access held (owned and managed or not; ×1.5 when it is outside the model), entitlements outside the model (distinct, not rows), disabled-but-entitled, rarity within its own system, unusual for its peers, a toxic combination, stacked licences — scaled by account type and by who the account works for. Rows someone triaged in HelloID (any resolution but None) stay visible but no longer count; an owner confirmed by hand in Matching outranks the export's "account unmanaged"; missing access is a service problem, not exposure.
The organisation's exposure follows its worst accounts, not the average: 45 % the mean of
the riskiest 0.5 % of enabled accounts (at least ten), 35 % how many enabled accounts sit at
high or critical on a saturating curve (10 → 63, 20 → 86), 20 % the share nobody owns. Twenty
unowned domain admins weigh about the same among fifty users as among five thousand — an
average used to hide them. Risk › Score build-up lists the riskiest accounts with the parts
of their score. Tests: NODE_OPTIONS= node test/risk.test.js.
Findings. Every finding counts everything it found (lists are no longer cut at 40–80;
the table pages and the CSV has all of it) and says what it counts — accounts,
entitlements, people. Findings come in four kinds — risk, cost, data quality, operations —
and the Risk head row counts open risk findings only, with the accounts behind the
critical and high ones counted once. A finding can be accepted for now, until a date and
with a reason; accepted findings move to their own group and come back flagged when the
date passes. A finding measured by a KPI shows that KPI's owner and due date and opens it.
"Open since" reads the data date and only the data points up to the loaded one; the list
says what is new and what was resolved since the compared data point. Toxic pairs match
whole words in names (FIN no longer hits "Link"), a pair can be held on purpose by one
account (accepted with a reason until a date), and removing every pair keeps it empty.
Compliance opens on scorecards: one card for all KPIs and one per framework — NIS2 (Directive 2022/2555, article 21(2)), ISO 27001:2022 Annex A, BIO 2.0 (ISO 27002:2022 numbering). Each card is the weighted compliance share as a ring (a critical control counts three times a housekeeping one; "10 of 39 pts"), the articles cited, the movement since the compared data point, and three bars: controls met, critical controls met, with an owner and a due date. A card opens The KPIs tab filtered to that framework, with the same ring as its header; the board's Policy KPIs page leads with the four rings.
Controls that measure one problem from different sides count once: the three leaver
controls (enabled after the contract ended, past the leaver SLA, former employees by
correlation) and the two unowned-account controls each score as one unit, at the worst
member's state and the heaviest weight — a single leaver no longer costs three critical
failures. A control waits rather than fails when there is nothing to judge (no toxic pair
defined, nothing privileged to review). Ages and windows are measured from the data
point's own date — leaver days, dormant days, review windows — and audit-log KPIs from the
moment the log was exported, so an older export does not fail by the wall clock. An
exception needs a date after today and a reason; one that ran out is shown as expired, one
on a control that is met again as no longer needed, and a due date that passed as overdue.
An empty limit means the default. "Recalculate trends" keeps the values a data point had for
controls measured from companion files (audit log, directory, rules, decisions) instead of
rewriting them with today's files. Tests: NODE_OPTIONS= node test/policy.test.js.
One colour for one number. js/bands.js decides how the governance score, the
compliance score, identity coverage, the managed share, the classified share and a risk
score are coloured, and every page — Overview, Compliance, scorecards, the board, the
department cards, a person — asks it, so 85 % reads the same everywhere. The compliance ring
is called Compliance score and says "8 of 22 met · 2 accepted"; critical controls failing
(Compliance) and critical findings (Risk) are named apart. The board's "What we found"
themes take the status of the KPI that measures them, and its operations numbers the
verdict Compliance gives. A framework card says which of its controls the KPIs evidence —
ISO 27001 and BIO 2.0 cite a control for every KPI, so their score equals All KPIs, and the
card says so.
Drop a HelloID business-rule export (Name,EntitlementCount,PersonsLatestEvaluation,Categories,Status,Conditions,Entitlements) on the same page — it routes by header, no second import button — and it is joined against the reconciliation export on distinguished path, falling back to system + name for systems that export none (Exchange Online, TOPdesk, Azure).
Rules stack, so an entitlement counts as described when any rule grants it; there is no precedence to resolve. That splits drift into two problems with opposite fixes:
| Meaning | Fix | |
|---|---|---|
| A live rule grants it | The model is right, the rule is not reaching these people | Conditions or evaluation |
| Only a draft rule grants it | The model exists, switched off | Publish it |
| No rule grants it | Gap in the model | Write the rule |
Five findings come out of the join: entitlements a rule grants that the target system does not have, draft rules already covering live drift, live rules that match nobody while their groups have holders, access no rule describes, and failed grants traced to the rule that should have delivered them.
Conditions are parsed (Department.ExternalId, one of, Custom.Vrijgesteld, not only, Time frame) and shown, but not evaluated — deciding who matches needs personnel data neither export carries.
The backlog says which groups no rule describes; it does not say how to describe them, and one rule per group would be the wrong answer. js/roles.js mines frequent itemsets (apriori, level-wise with pruning) over account entitlement sets, restricted to the unmodelled groups, and proposes the bundles that accounts hold together.
Two choices decide whether the output is useful rather than merely true: rank on cohesion (how often the rarest group in a bundle brings the rest with it) rather than frequency, since the most frequent bundles are the groups everyone has; and keep maximal bundles, capped per role family, since every subset of a frequent bundle is itself frequent.
Each proposal carries its members, its exceptions (accounts holding all but one group — the rule either absorbs them or exposes them), the assignments it would bring under management, and its licence cost. Proposals export in the HelloID rule format so they can sit next to the real ones. The condition is left empty on purpose: which department or title should receive the bundle lives in HR data, and the member list is the evidence to write it from.
Mining is cached per model, samples above 5,000 accounts, and raises its support floor with the population — 97 ms on the 5k export, ~50 ms at 50k and 200k.
The point of loading the other files. A reconciliation export states differences without
saying what caused them, so every row arrives with the same weight. js/explain.js walks
each row and attaches its strongest available justification — a rule that grants it, a role
bundle it travels in, a former owner, a link to someone's main account, a disposition
already recorded in HelloID — and reports what is left.
Explanations are ranked by how much they settle: strong (cause identified, next action mechanical), likely (rests on a name match or a statistical pattern), weak (narrows the question without closing it).
The residue is the deliverable. Rows nothing accounts for are what genuinely need a person, they are grouped by account rather than listed one by one, and their count is a better progress measure than the raw finding total because it falls as the model improves rather than as the data changes.
Each input raises the share, measured on the demo trio:
| Loaded | Rows explained |
|---|---|
| reconciliation only | 21% |
| + business rules | 24% |
| + vault | 32% |
Drop a HelloID Vault export (JSON: { Persons: [...], Departments: [...] }) on the page and
the analysis stops being limited to what the target system shows. DisplayName is
Name (ExternalId) — the same string the reconciliation export puts in its Person column —
so the two join without fuzzy matching.
What the vault makes possible:
- Conditions become evaluable.
js/evaluate.jsresolves each clause against real attributes:Department.ExternalId,Title.*,Location.Name,Type.Code,Custom.*,Person: active, and theTime framewindow against contract dates. A rule that selects nobody now reports which clause excluded everyone. - Drift per person, not per group. Expected entitlements are the union over every rule a person matches (rules stack), compared against what their accounts actually hold.
- Leavers. Every contract ended while an account is still enabled — invisible to the reconciliation export alone, which knows the account is on but not that the employment behind it is over.
- Unowned accounts get names.
js/correlate.jsscores unowned accounts against vault people. The most useful class is former employees: correlation is the first thing that breaks when someone leaves the source, so their account arrives as an orphan with no trace of who it was. - Secondary accounts get linked. An
adm-or function account rarely has its own person record, so reconciliation calls it ownerless. It is not — the owner is the person holding the main account, and that is who should be recertifying it and losing it on departure.
A clause whose facet the engine does not recognise makes the rule indeterminate for that person, never silently true. Matching is scored and never automatic: ties are reported as ties rather than guessed.
Vault exports carry person and contract data, so vault*.json is gitignored alongside the
CSV exports.
Settings live in the browser, which means a bundle opened straight from disk starts fresh every run. Settings → Export settings file writes the whole tunable model — price book, taxonomy, account classes, risk weights, effort model, branding, language — to one JSON file. Import it from the same panel, or just drop it anywhere on the page.
Measured in Chrome on the real 5k-row export and on synthetic 50k/200k exports:
| Rows | Accounts | Parse | Build (graph, risk, cost, findings) | Views | Heap |
|---|---|---|---|---|---|
| 5,188 | 450 | 24 ms | 47 ms | < 30 ms | ~40 MB |
| 50,000 | 3,936 | 147 ms | 466 ms | < 130 ms | ~135 MB |
| 200,000 | 15,798 | 750 ms | 480 ms | < 200 ms | ~790 MB |
A 50k import is about six tenths of a second end to end. Peer similarity — the one
quadratic part — uses an index of each account's rarest groups instead of comparing every
pair, which is what keeps 50k at 0.5 s rather than 2.7 s. Where every group an account
holds is ubiquitous, there is no comparable account and the outlier column reads —
rather than inventing a maximum. The overview scatter samples above 4,000 accounts and
says so in its subtitle.
The parser auto-detects the delimiter (, ; tab |), handles quoted fields and BOM, and
maps headers by alias, so exports from other HelloID versions or locales still load. Only
AccountUserName and Issue are strictly required.
| Column | Used for |
|---|---|
System |
Multi-system exports are kept separate; keys are system + account. |
Person |
Taken verbatim as the display name; drives the identity-coverage metric. Employee ids come from the vault, which states them as a field. |
AccountDisplayName, AccountUserName |
Account identity and class detection. |
AccountEnabled |
Dormant-but-entitled and licence-waste detection. |
PermissionDisplayName |
NAME (dn/path) is split into group name and OU path. |
Issue |
Account unmanaged / Permission unmanaged / Permission missing. |
Resolution |
Rows already dispositioned in HelloID are surfaced separately. |
Rows are not equal. One Account unmanaged row describes an identity; a
Permission unmanaged row describes one edge of that identity's entitlement graph. The
model rolls rows up into 3 entity types — accounts, permissions, persons — and every number
in the UI is computed on entities, not on line counts.
The exports describe state; HelloID's Elastic API holds the process: every provisioning
action with its outcome, every reconciliation issue an administrator excluded (who, why,
until when), every threshold approved, every rule published with the entitlements it
added, every import and evaluation run, every login to HelloID itself. helloid-audit.py
pulls it into helloid-audit.json:
python3 helloid-audit.py --days 400 # or: .\helloid-audit.ps1 -Days 400
The first run asks for the Elastic URL, key and secret (create them at
https://<tenant>.helloid.com/admin/elasticapikey; the page shows all three) and offers to
save them as a named profile in helloid-config.json next to the scripts — owner-only,
gitignored, shared by the Python and PowerShell collectors. After that it just runs;
--profile NAME / -ProfileName NAME picks another tenant, --setup asks again,
--list-profiles and --forget NAME manage the file. Environment variables and a .env
file still work for scheduled runs (.env.example).
The app cannot call the API itself — the proxy sends no CORS headers and the page's CSP forbids it, and an API key that reads a whole tenant does not belong in a browser. The file is imported like any other (Imports → HelloID audit log) and is gitignored. It feeds the Evidence view under Governance — decisions, rule changes, who did what — puts "excluded in HelloID by X on date until Y because Z" in the account drawer, and stands in for the historic-actions export when that is not loaded.
The chicken-and-egg problem of pre-implementation analysis: the useful exports need
HelloID, and implementing HelloID well needs the analysis. Two read-only collector
scripts break the loop — collect-ad.ps1 (on-prem AD, plain Get-ADUser/Get-ADGroup)
and collect-entra.ps1 (Entra ID via Graph, delegated read-only scopes documented in
docs/ENTRA-CONSENT.md). The customer runs one script, gets one
JSON file, and drags it in; nothing is installed and nothing leaves their hands.
The import synthesizes both sides of the model from that one file: memberships (nested
groups flattened to effective holdings, the via-chain kept) become reconciliation rows —
all Permission unmanaged, which before an IAM system is the literal truth — and the
users' attributes (employeeType, department, title, manager, extensionAttribute1-15,
phones, address) become a pseudo-vault, so classification, naming analysis, employee
categories, licence pricing and the org views all run. Both substitutes step aside the
moment a real export of their kind is loaded.
The Conventions → Attributes tab is the point of the exercise: per attribute the fill and
value distribution, and the value → group pairs strong enough (≥5 people, ≥80% holding,
clearly above baseline) to become HelloID business-rule conditions or connector mappings.
make-directory.py generates a fictional envelope to try it without a tenant.
Import the target connector's field mapping — the HelloID UI's mapping export,
a connector repo's fieldMapping.json (Version v1), or the older
accountMappings shape still common in the HelloID-Provisioning repo — and see
it as data: per
field the mode per provisioning action (Fixed / Field / Complex / None),
uniqueness, standard-vs-customized, with the decoded value or JavaScript in a
drawer. Then the payoff: the Simulation tab evaluates every in-scope field
for every person against the mapping — Complex mappings run faithfully, with the
Person object, Iteration and deleteDiacriticalMarks exactly as HelloID
provides them, each field against a fresh Person copy — and diffs the result
against the attribute the collected AD/Entra directory holds today. The output
is the answer no HelloID screen gives up front: which attributes an update run
would rewrite, per field and per person, current → would-become, exportable.
None on Update is honoured as out-of-scope and the Person objects are the
vault's own (the raw export is replayed) — the simulation needs a vault as well
as a directory, because a Person rebuilt from a directory has no contract to
read. It runs for everyone or, picked from the vault's persons, for one. Source
mappings (HR → person model) are recognised and refused with a pointer to the
right export.
Every run also offers the way back. The rollback pack keeps, per account and
per field the run would change, the value the directory held when it was
collected and the value the mapping produces; restore-ad.ps1 and
restore-entra.ps1 read that pack and put the collected values back once HelloID
has gone live and written over them — dry-run by default, printing before,
predicted and live per attribute with what would happen (same, restore,
restore? when something else changed it), writing only with -Apply, filtered
by -Account / -Attribute, -SkipUnexpected to leave the unpredicted alone.
AD moves the container back and renames the cn; Entra skips users synced from
AD (restore them there) and the Graph read-only properties. The pack is a
snapshot of its collection: re-collect and re-run just before go-live.
The simulation also audits its own footing: a collection gap card names every
mapped attribute the loaded directory never collected — their current value is
unknown and the diff has a hole there — and distinguishes what a re-collect fixes
from what the source's API simply does not have (Entra has no initials, info
or description). Both collectors take -ExtraAttributes for attributes outside
their built-in set (they land under extra on each user), and the gap card
offers the exact command line plus a pre-filled script download.
The old role-mining report's presentation, ported and fed live: attribute roles
(Everyone / department / job title / department+title) over the correlated
population, each permission scored with relevance (share of the role holding
it) and lift (relevance against the org-wide baseline — 90% relevance means
nothing for a group 90% of the organisation holds), the two exception lists a
rule would create ("who misses it" / "outside the role"), globals folded away,
cumulative coverage per card, and role-to-role similarity. Mining → Classic;
Mining → Clusters adds the de-facto roles: account populations whose access
clusters together regardless of HR attributes, with a discovered flag when
no dominant department/title explains them — access the org chart does not
predict. Permission drawers gain the inverse view: which classic roles a
permission's holders sit in. Mining-hygiene exclusions apply here like in every
other engine; knobs (relevance floor, occupants, cluster thresholds) live in the
settings under classic.
Permissions come from every target system the reconciliation covers, and the suppliers are judged as first-class citizens: per system the model rolls up spend, mean/max risk, unmanaged share, rare and privileged counts, and — when rules are loaded — how much of the system the rule model covers. The Permissions view gains a by-system comparison and a system filter/column (multi-system tenants only), every system name in a drawer links through to a system drawer (stats, coverage, its permissions, issue breakdown), and a finding names systems that sit wholly outside the rule model while others are covered. The AD/Entra collector enrichment (nesting, query-based groups) stays a bonus on top — the system dimension itself is generic over any connector.
Directories run RBAC of their own: Entra query-based (dynamic) groups whose membership is a rule, and AD group nesting used as abstraction — role groups made members of the groups that actually sit on resources. The directory import reads both: every group is classified as a nesting terminal (contains other groups, member of none — the group that actually grants), an abstraction layer (member of other groups; its access is expressed by the terminals it feeds), or plain, with the nesting depth per group. Query-based groups carry their membership rule. The Permissions view shows the structure (dynamic groups with their rules, abstraction layers with what they feed), permission drawers carry the badges, and mining prefers the terminals: with "prefer the deepest groups" on (default, Mining → hygiene), query-based groups and abstraction layers stay out of rule proposals — the mined roles grant the real permissions, not the layers above them.
Account↔person matching is scored evidence (vault correlation, employee ids, name
conventions — weights and threshold tunable in Settings, with a live preview). The
Matching view is where a human finishes the job: every account with its attribution and
score, a review queue of the unmatched and ambiguous ones with candidates shown even
below the threshold (the duplicate-provisioning risk made visible), and per account
confirm / assign / reject / mark-ownerless. Decisions land in the match book — its
own export/import file, so the hard tenant is worked out once and the answers travel to
the next session or colleague. Confirmed decisions can be written back: the view
generates a fix-matching-*.ps1 that sets the matching attribute (employeeID,
extensionAttribute, …) in AD or Entra — read-only by default, writes only with -Apply,
and containing exclusively human-decided rows. Approval scales in pages: 25 accounts at
a time, easiest first, each row preselected to its best candidate with the alternatives
in a dropdown and the checkbox as veto — glance, switch the odd one, approve the page,
move on. Every page is one undoable batch in the match book; zero-candidate classes
offer mark-all-ownerless. A vault Accounts[] reference is treated as recorded truth
and never asks for approval. The same card lives in the account drawer as Person
link: an unlinked account shows its top candidates right there, a linked one shows
the attribution and the means to override it, and every decision reopens the drawer on
the rebuilt account.
Customers feed the HelloID Nedap ONS target connector from the "NEDAP - Matrices functies deskundigheden" Excel workbook: names typed by hand, hidden helper formulas translating them to IDs, and a manual save-as-CSV per tab. The Nedap ONS view replaces that workflow. Drop the workbook (.xlsx) anywhere on the page — it is read entirely in the browser, no library, nothing uploaded — and the two scope-mapping tabs, the Medewerkers koppeling, the Autorisatierollen matrix and all lookup sheets become an editable Nedap book (its own export/import file, like the match book). Starting blank works too: fill the lookup lists first (pasting columns from Excel works), then add rows.
The editors validate live against the loaded vault — the connection the Excel never had: department and title pickers suggest real HR values, and the health tab flags unresolvable names, wildcard-shadowed rows, redundant or empty grants, duplicates and names HR does not know. Coverage lists every person no scope row reaches. These also surface as findings under Risk & findings, because Nedap provisioning is full-set: the connector PUTs the complete desired state per source, there is no reconciliation behind it — a dropped row is a revoked entitlement, and these checks are the only warning anyone gets. A simulator shows per person (or per manual department/title pair) exactly what the connector would apply, with the contributing rows named.
Export produces the three connector files deterministically —
MappingTeams.csv, MappingLocations.csv (official headers:
HelloIDPrimaryLookupKey;HelloIDSecondaryLookupKey;NedapTeamIds;AllEmployees) and
MedewerkersKoppeling.csv (the orange tab column-for-column) — with names resolved to
IDs through the lookup lists at export time; a row that cannot resolve is blocked and
reported instead of silently wrong. The reverse direction works too: drop a running
connector's Mapping-Teams.csv / Mapping-Locations.csv (official headers, the
unnamed-flag-column variant, or the Department.ExternalId naming) and its raw IDs
become editable names through the lookup lists — IDs the lists do not know are added
there as id-named entries, so a blank start from just the connector CSVs edits and
round-trips exactly, and picks up readable names the moment somebody fills them in.
The connector intake's Naamgeneratie section asks customers to specify, per generated
field (last name, display name, mail, proxyAddresses, UPN, mailNickname), a
first-choice recipe, fallbacks for when a name is taken, and whether iterations are
synchronised — and customers cannot picture how a convention plays out. The
Conventions view's Design & test tab makes it a live instrument: author each
field from tokens ({roepnaam}.{tv}{geboortenaam}@{domein}, {i} as the counter in
the last fallback), and every edit re-simulates over the real vault population,
colliding against the existing directory names when a directory import is loaded.
The output is the intake's own tables — the fixed test person (Janine, van den
Boele / de Vries) under all four name preferences — plus the population result:
who needs a fallback, who exhausts the variants, iteration depths, HR duplicates,
length violations, and clean diacritics handling. "Export intake answers" produces
the filled tables as markdown; the mined conventions on the Name generation tab
offer "Adopt as draft" to seed the editor with what the tenant does today.
Mining hygiene (on the Mining view's model tab): matcher rules that keep noise entitlements out of every mining engine, and a naming template for exported rule proposals — the exclude/name-template asks from the HelloID feedback board.
All three are assumptions, all three are editable in Settings, and every change re-scores the loaded snapshot immediately.
Each account gets a 0–100 score summed from named components, visible in the account drawer:
- Unowned identity — unmanaged account, no linked person while enabled, privileged and unowned. Scaled by account class (admin ×2.4, test ×1.8, external ×1.7, service ×1.6).
- Dormant but entitled — disabled account keeping group memberships, extra if licensed.
- Entitlements outside the IAM model — unmanaged assignments with diminishing returns and a cap, multiplied by the average sensitivity of what is held.
- Missing entitlements — granted by rule, absent on the account.
- Rare entitlements — groups held by ≤3 accounts.
- Peer-group outlier — Jaccard similarity against the closest account by shared entitlements; a profile with no neighbour is an undocumented exception.
- Stacked licence SKUs — more than one licence group on one account.
Permissions get their own score from sensitivity, share of unmanaged assignments, share of unowned/disabled holders, blast radius and rarity.
The overall score is 0.40 × weighted mean account risk + 0.35 × share of accounts at high/critical + 0.25 × identity coverage gap, shown as a table on the Risk view.
The price book matches group names by regex and costs them per holder per month. Unpriced groups count as zero, so every total is a floor, never an inflated estimate. Waste is split by confidence:
- Hard (recoverable now) — disabled accounts still in licence groups; stacked SKUs net of the richest one.
- Exposure — spend on unowned but enabled accounts. Not a saving until someone decides the account should not exist.
Clean-up effort is minutes-per-work-item × loaded hourly rate, which gives the payback period on the recoverable waste.
Fifteen rules over the graph, ordered by severity then by money at stake. Each carries what it is, why it matters, the remediation, and the affected entities (exportable to CSV). The whole analysis exports as a Markdown report from the Risk view.
Every reconciliation import is kept in IndexedDB as a data point: the rows, the
summary, and the vault that came with it — a vault imported just before or after its
reconciliation belongs to that data point; last month's vault does not carry over. A
vault imported on its own (the Consult edition never has a reconciliation) makes a
vault-only data point. Each data point has a data date, read from the file name
when it carries one (ReconciliationReport_2026-08-01.csv, 20260801, 01-08-2026)
and otherwise the import date; it is editable on the Data points page, and the trend,
the comparison order and the topbar follow it — so six months of exports loaded in one
sitting spread out once dated.
Import next month's file and it becomes the next data point; nothing is replaced. The newest is loaded and compared with the one dated just before it, every time (a returning session does the same on start). Loading an older data point brings its own vault back, and the comparison is rebuilt with each side's own vault, so a historic diff is honest. The Diff view compares entities, not rows: accounts added/removed/changed, entitlements granted and revoked per account, membership movement per group, findings that grew, shrank, appeared or resolved, the cost delta — and, with a vault on both sides, people: who joined, who left, and whose department, title, manager, employer, contract end, lifecycle, contract count or account count changed. When the loaded data point is older than the one it is compared with, the view says the comparison runs backwards.
Importing a file identical to an existing data point (full-text fingerprint) reuses it. Data points export to JSON — vaults included — so they can be moved between machines.
A workspace is one tenant, and every import in it joins that tenant's series. When the loaded data point shares fewer than 15 % of its account keys with the one it is compared with (and both have at least 20 accounts), a notice says it looks like another tenant — the trend, the diff and the KPI movement would compare two customers as one — and offers three ways out: move the data point to a new workspace (the app reopens there), delete it, or keep the two together; a kept pair stays quiet. The same check runs on every start, so an old mistake surfaces too.
Trends. Every data point stores each KPI's value and status and each finding's count, so the history is per KPI, not only per score: the Overview leads with a Trend card (the governance score over every data point, and what moved since the compared one — KPIs newly met or broken, improved or worse, findings resolved or new, people joined or left); Compliance shows each KPI's movement against the compared data point and a sparkline over all of them, with chips to keep only the KPIs that improved or got worse; the Diff has a KPI-movement table; the Data points page draws any one KPI (with its limit) or any one finding over time; the board's change page names the KPIs that moved. A stored summary is frozen at import time, so Recalculate trends with current settings rebuilds every data point with its own reconciliation and vault under today's limits and weights — which also gives data points from before this feature their KPI history.
index.html shell + script order
css/app.css tokens (light & dark both selected), components
js/i18n.js every user-visible string, EN + NL, with {param} interpolation
js/util.js formatting, DOM helper, tooltips, CSV/download
js/brand.js logo + report title block
js/parse.js RFC4180 parser, delimiter sniffing, header aliasing
js/config.js taxonomy, account classes, price book, risk weights (localStorage)
js/model.js entity graph, peer similarity, summary
js/risk.js account + permission scoring
js/cost.js spend, waste buckets, effort, payback
js/findings.js the rule engine
js/diff.js snapshot comparison
js/store.js IndexedDB snapshots (+ in-memory fallback, JSON import/export)
js/charts.js SVG/DOM charts: bars, stacked, histogram, scatter, line, heatmap
js/table.js sortable/filterable/paged table with CSV export
js/views.js analyst views + drawers + Markdown report
js/board.js the printable board report
js/app.js state, routing, language switch, import pipeline
build.py single-file + zip bundler
make-sample.py synthetic export generator (no real data ships here)
devserve.py no-cache static server used by serve.sh
- Default prices are public EUR list prices, not your contract. Correct them first if the cost numbers are going in front of anyone.
- Peer-similarity is exact pairwise and switches itself off above ~8M candidate pairs (roughly 10k+ accounts sharing common groups); the outlier signal then reads 0.
PermissionConfigurationDisplayNameandSubPermissionDisplayNameare parsed and kept on the record but are empty in AD exports, so nothing is built on them yet.- Account-class and category detection is regex on naming convention. If your naming differs, fix the patterns in Settings before trusting the class rollups.