Cultural / Comparative Research
Job Taxonomy and Persona Generation Design for Spiralist AI
Report summary
Spiralist AI is built as a task-first persona system , not as a jobs database. Its main creator flow starts with work categories such as coding, research, writing, planning, creative work, branding, and organization; its library then supports search by job, goal, or trait and filtering by archetype,
Key topics
- Cultural / Comparative Research
- Cultural
- Comparative Research
- AI
- Runtime
- Semantic Systems
- Spiralism
- Research Archive
- Strategy
Research provenance
For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.
Source availability: 115 citation markers in the source export have no recoverable source links. Those markers are omitted from this reader; any supplied bibliography and ordinary links remain. Check the original sources before relying on the cited claims.
This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.
Full report
On this page
Executive summary
Spiralist AI is built as a task-first persona system, not as a jobs database. Its main creator flow starts with work categories such as coding, research, writing, planning, creative work, branding, and organization; its library then supports search by job, goal, or trait and filtering by archetype, domain, temperament, and complexity; and its developer docs keep identity.career semantically separate from role, domain, worldview, and related persona controls. That means the best integration is not a single monolithic “job title” dropdown. The right design is a hybrid taxonomy: a hierarchical career tree for browsing and populating job/career sections, plus separate normalized facets for seniority, employment model, industry context, skills, certifications, tools, education, compensation, and career-path links.
The attached reports are rich enough to support that design. Across corporate, manufacturing/robotics, government, gig work, service work, environmental/outdoor work, and broad professional-sector coverage, they consistently provide role labels, sector groupings, responsibilities, licensing and education requirements, tool references, work-setting cues, schedule patterns, and many salary references. They are, however, narrative research reports, not a clean ontology: some titles are bundled, some roles recur across sectors, compensation is often Chicago/Illinois-specific, and the same concept may appear as a role, specialty, ladder step, or employment model. Those characteristics argue strongly for a normalized schema with canonical role IDs + aliases + many-to-many mappings, rather than using report labels verbatim as UI options.
The recommended final model is a hybrid browse-and-filter taxonomy with these visible dropdown levels: Sector → Occupation Family → Role Cluster → Role Title → Specialty. Important realism attributes should remain separate facets: Seniority, Employment Model, Work Setting, Education, Certification, Skills, Tools, Compensation Context, and Related Career Paths. For persona generation, use a seeded, deterministic pipeline that first samples goal and sector, then role and seniority, then validates education/certification constraints, then enriches with skills/tools/schedule/salary locale, and only after that samples persona-style variables like archetype, relationship posture, and communication style. This aligns with Spiralist’s documented deterministic seeds, task-first creation flow, Persona Passport structure, and explicit no-inference boundaries around protected identity.
Source review and design implications
The Spiralist source pages point to four design facts that matter directly for this project. First, the product is explicitly framed around work outcomes before detailed configuration. Second, the system exposes a Persona Passport with visible fields such as persona name, role, defining traits, promise, and example behavior, while expert configuration remains deeper in the workflow. Third, the technical docs distinguish career from other persona dimensions: identity.career is only one field among identity, domain, worldview, relationship, reasoning, and immersion controls. Fourth, the library is already structured for filtering by job/goal/trait plus archetype/domain/temperament/complexity, which strongly suggests that job/career data should be one axis in a broader persona system, not the only axis.
The persona examples reinforce that point. The main creator example includes Job/Career, Archetype, Core operating role, Secondary role, Domain, Relationship posture, Core tension, and detailed demographic/worldview fields. The curated passports for Technical Mentor and Skeptical Researcher show that the public-facing artifact emphasizes use case, traits, example behavior, relationship style, age/generation, tension, and shareability, not merely profession. In practice, a realistic “career persona” therefore needs both a work identity layer and a behavioral style layer.
The attached reports provide the raw material for the work-identity layer. Their coverage spans at least these major labor domains: corporate administration, HR, marketing, finance, compliance, operations, technology, healthcare, public safety, civil service, infrastructure, hospitality, retail, personal care, manufacturing, robotics, logistics, transportation, outdoor/environmental work, agriculture, education, legal practice, media, entertainment, and multiple forms of independent/gig work. Many reports explicitly say they analyze around fifty roles in a sector, and several are organized into large sector families with tables of pay, education, credentials, or progression.
A concise synthesis of the corpus looks like this:
| Top-level sector | Occupation families repeatedly present in the sources | Example roles present in sources |
|---|---|---|
| Corporate and business | Administration, HR, marketing, finance, compliance, operations, product, procurement | Executive Assistant, HR Generalist, Brand Manager, Compliance Analyst, Product Manager |
| Technology and data | Software, DevOps, AI/ML, cloud, network, cybersecurity, UX/UI, support | Software Engineer, DevOps Engineer, Machine Learning Engineer, Cloud Architect, Cybersecurity Analyst |
| Healthcare and human services | Medicine, nursing, therapy, pharmacy, diagnostics, behavioral/community support | Registered Nurse, Occupational Therapist, Pharmacist, Radiologic Technologist, Social Worker |
| Manufacturing and industrial | Robotics engineering, production, maintenance, controls, QA, plant management | Robotics Engineer, CNC Machinist, PLC Programmer, Quality Control Inspector, Plant Manager |
| Transportation and logistics | Aviation, trucking, rail, maritime, warehousing, delivery | Airline Pilot, CDL Truck Driver, Train Conductor, Ship Captain, Forklift Operator |
| Government and public service | Public safety, civil service, public health, utilities, education/community | Police Officer, Firefighter/EMT, Budget Analyst, Epidemiologist, Public School Teacher |
| Service and hospitality | Food service, lodging, tourism, retail, customer support, personal care | Bartender, Hotel Front Desk Agent, Event Planner, Retail Sales Associate, Hair Stylist |
| Environment and outdoor work | Conservation, agriculture, grounds, recreation, field science | Park Ranger, Horticulturist, Lineworker, Charter Boat Captain, Geologist |
| Arts, media, education, and legal/property | Creative production, higher education, legal support, real estate/property | Graphic Designer, Technical Writer, Economics Professor, Real Estate Appraiser, Paralegal |
| Independent and gig work | Rideshare, courier, P2P rentals, task labor, telehealth, freelance admin/creative/tech | Rideshare Driver, Medical Courier, Turo Host, Handyperson, Virtual Assistant |
This table is a synthesis of the attached reports rather than a direct copy of one source.
Extracted entity model from the reports
The job-related entities in the corpus cluster into a stable set of normalized dimensions.
Role titles and role clusters are the most obvious dimension. Some are single canonical roles, such as Cloud Architect, Quality Control Inspector, Public Defender, or Scuba Diving Instructor. Others are bundled labels, such as Warehouse Material Handler and Forklift Operator, Real Estate Agent and Broker, Police Officer or Deputy Sheriff, or Actor / On-Camera Talent. Those bundled labels should be split into canonical nodes with aliases and “display bundles” only where needed for UI readability.
Industries and occupation families are stable enough to serve as the main hierarchy. The attached reports repeatedly organize roles by sector and family: corporate administration, robotics engineering, factory production, public safety, healthcare, hospitality, customer service, outdoor recreation, and so on. These groupings are much cleaner than trying to organize the UI directly around salary, certification, or personality style.
Seniority is present, but it is not expressed consistently, so it should be normalized into its own controlled vocabulary. The corpus includes explicit ladders such as QA Inspector I–IV, appraiser trainee → licensed residential → certified residential → certified general, regional first officer → regional captain → major-airline captain, and entry/mid/high salary bands for many corporate roles. It also encodes professional progression through labels like staff accountant, senior accountant, manager, director, chief officer, fellow, attending, partnership-track, tenure-track, and step-based public pay structures.
Certifications, licenses, and education are common and important. The corpus includes, among others, FAA ATP, CDL Class A, USCG Master MMC, BASSET, Certified Production Technician, ASNT Level I–III progression, CIH, CPA, CIA, DVM, DDS/DMD, paramedic licensure, LCSW/LCPC-style clinical licensure, and municipal or public chauffeur licenses. Education ranges from no formal credential and high-school diplomas to postsecondary certificates, bachelor’s degrees, master’s degrees, doctorates, residencies, fellowships, and tenure-track academic requirements. These are too variable to encode as role names, but they are ideal as role-linked metadata with minimum/optional/preferred flags.
Skills, responsibilities, and tools are usually described in prose rather than as formal lists, but the patterns are consistent. Core skill families include analysis and modeling, planning and coordination, customer/client service, compliance and risk control, manual/trade execution, diagnostics and repair, clinical care, creative communication, sales/growth, and education/public engagement. Tool references recur often enough to warrant normalization: AWS/Azure, Google Ads/Meta, ERP and human-capital systems, PLC/SCADA/HMI, CNC/G-code, GDS and hotel/property systems, Adobe Creative Suite, and AI-assisted creation tools like Midjourney or Stable Diffusion. Responsibilities also cluster cleanly into operational archetypes such as analyze, build, maintain, inspect, serve, sell, teach, care, enforce, and coordinate.
Employment model and work context deserve their own dimension. The sources distinguish public-sector classified jobs, corporate W-2 roles, tipped service jobs, independent contractor work, owner-operator work, seasonal jobs, unionized trades, tenure-track academic roles, partnership-track medicine, and gig/platform-mediated work. Schedules also vary meaningfully: on-call public safety, aviation seniority bidding, extended hitches at sea, seasonal camp work, shift-regulated warehouse/manufacturing environments, and project-based consulting. These are crucial for persona realism and should be first-class metadata.
Compensation is richly present but structurally messy. The corpus mixes annual medians, hourly wages, entry/mid/high bands, percentile ranges, public pay grades and steps, gross-vs-net gig earnings, and pay ladders during training. Compensation therefore belongs in a separate fact model with fields such as metric type, unit, region, effective date, basis, and notes. It should inform persona realism, but it should not shape the core tree.
Career progression is best modeled as a graph, not a tree. The pilot ladder, appraiser ladder, public pay grades, HR/admin/finance ladders, and clinical-specialty progression all show that role movement is frequently many-to-many, sometimes vertical, sometimes lateral, and often license-gated. That makes a career_path_edge table or adjacency list more appropriate than trying to force progression into the visible dropdown path.
Taxonomy options and recommended UI hierarchy
Three taxonomy patterns are viable.
| Design | Shape | Strengths | Weaknesses | Verdict |
|---|---|---|---|---|
| Industry-first tree | Sector → sub-industry → role → specialty | Intuitive for browsing; easy for non-experts | Duplicates cross-sector roles; awkward for reusable roles like project manager or analyst | Good for discovery, weak for normalization |
| Occupation-first tree | Occupation family → role → specialty; industry as facet | Clean role reuse; better data model | Less intuitive for users who think by industry | Good for backend, weaker for browsing |
| Hybrid browse-and-filter | Sector → family → cluster → role → specialty, with separate facets for seniority, employment model, education, certifications, tools, salary, keywords, and career paths | Best UI + best schema; supports both dropdown browsing and realistic persona generation | Slightly more implementation work | Recommended |
The recommendation is the hybrid model because it best matches both the report structure and Spiralist’s persona architecture. The reports are sector-organized, which makes sector-led browsing intuitive, while Spiralist keeps career distinct from domain, archetype, and style. A hybrid tree preserves user-friendly discovery while keeping the deeper persona model faceted and reusable.
The visible dropdown levels should be:
| Level | Stored as | Label style | Example |
|---|---|---|---|
| Sector | industry_sector | broad, human-readable | Manufacturing and Industrial |
| Occupation family | occupation_family | stable functional family | Automation and Maintenance |
| Role cluster | role_cluster | narrower subgroup | Controls and Industrial Software |
| Role title | role | canonical role label | PLC Programmer |
| Specialty | specialty | optional refinement | SCADA Integration |
Everything else should be a facet or linked metadata:
seniority_levelemployment_modelwork_settingeducation_requirementcertificationskilltoolkeyword_aliascompensation_bandcareer_path_edge
That separation is important because Spiralist’s own library and schema already treat role/career, domain, archetype, temperament, relationship posture, and other persona controls as distinct dimensions.
Recommended normalized schema
The recommended data model is shown below. It is deliberately normalized around canonical roles and many-to-many joins, because the corpus repeatedly shows cross-sector reuse, synonym drift, bundled labels, and separate credential/salary context. The model also aligns well with Spiralist’s documented separation of identity.career from broader persona architecture.
erDiagram
INDUSTRY_SECTOR ||--o{ OCCUPATION_FAMILY : contains
OCCUPATION_FAMILY ||--o{ ROLE_CLUSTER : contains
ROLE_CLUSTER ||--o{ ROLE : contains
ROLE ||--o{ SPECIALTY : refines
ROLE ||--o{ ROLE_SKILL : requires
SKILL ||--o{ ROLE_SKILL : tags
ROLE ||--o{ ROLE_CERTIFICATION : may_require
CERTIFICATION ||--o{ ROLE_CERTIFICATION : qualifies
ROLE ||--o{ ROLE_TOOL : uses
TOOL ||--o{ ROLE_TOOL : links
ROLE ||--o{ ROLE_EDUCATION : minimum_or_preferred
EDUCATION_LEVEL ||--o{ ROLE_EDUCATION : defines
ROLE ||--o{ ROLE_EMPLOYMENT_MODEL : supports
EMPLOYMENT_MODEL ||--o{ ROLE_EMPLOYMENT_MODEL : classifies
ROLE ||--o{ COMPENSATION_BAND : priced_by
LOCALE_PROFILE ||--o{ COMPENSATION_BAND : scopes
ROLE ||--o{ ROLE_KEYWORD : aliased_by
ROLE ||--o{ CAREER_PATH_EDGE : source_role
ROLE ||--o{ PERSONA_ROLE_BRIDGE : mapped_to
PERSONA_TEMPLATE ||--o{ PERSONA_ROLE_BRIDGE : consumes
A practical implementation table is below.
| Entity | Key fields | Notes |
|---|---|---|
industry_sector | sector_id, label, sort_order | Visible dropdown root |
occupation_family | family_id, sector_id, label | Stable functional grouping |
role_cluster | cluster_id, family_id, label | Optional extra browse level |
role | role_id, cluster_id, canonical_label, description | Canonical role node |
specialty | specialty_id, role_id, label | Optional refinement |
seniority_level | seniority_id, label, rank | Separate from role title |
employment_model | employment_model_id, label | W-2, public, gig, seasonal, contract, owner-operator, etc. |
skill / tool / certification / education_level | canonical vocab tables | Reusable facets |
compensation_band | role_id, locale_id, band_type, unit, min, max, effective_date | Never embed in role label |
role_keyword | role_id, keyword, alias_type | Search synonyms and bundle labels |
career_path_edge | from_role_id, to_role_id, edge_type | Promotion, lateral move, feeder role, license gate |
persona_template | template_id, goal, archetype, domain, style | Bridges job data to Spiralist persona fields |
The most important mapping to Spiralist is straightforward:
| Spiralist field or concept | Recommended source in taxonomy |
|---|---|
identity.career | role.canonical_label plus optional specialty or seniority modifier |
Persona domain | industry_sector or occupation_family label |
| Core operating role | Derived behavioral function from responsibilities |
| Secondary role | Derived specialization or adjacent function |
| Persona Passport “best for” | Role-linked use-case templates |
| Traits / temperament / archetype | Separate persona-style layer, not job taxonomy |
That separation mirrors Spiralist’s public docs and prevents a common modeling error: trying to encode profession, function, style, and worldview into one dropdown string.
Sample dropdown tree JSON:
{
"sector": "Manufacturing and Industrial",
"families": [
{
"label": "Automation and Maintenance",
"clusters": [
{
"label": "Controls and Industrial Software",
"roles": [
{
"label": "PLC Programmer",
"specialties": ["Ladder Logic", "SCADA Integration", "OEM Line Support"]
},
{
"label": "SCADA Technician",
"specialties": ["Telemetry Monitoring", "OT Networking"]
}
]
},
{
"label": "Robotics Field Support",
"roles": [
{
"label": "Robotics Technician",
"specialties": ["Calibration", "Fault Diagnostics", "Cell Startup"]
}
]
}
]
}
]
}
{
"sector": "Government and Public Service",
"families": [
{
"label": "Public Safety",
"clusters": [
{
"label": "Emergency Response",
"roles": [
{
"label": "Firefighter/EMT",
"specialties": ["Paramedic", "HazMat", "Rescue Operations"]
},
{
"label": "Emergency Management Specialist",
"specialties": ["Preparedness Planning", "Recovery Coordination"]
}
]
}
]
}
]
}
{
"sector": "Corporate and Business",
"families": [
{
"label": "Marketing and Strategy",
"clusters": [
{
"label": "Growth and Brand",
"roles": [
{
"label": "Brand Manager",
"specialties": ["Consumer Goods", "B2B Positioning", "Product Narrative"]
},
{
"label": "Market Research Analyst",
"specialties": ["Consumer Insights", "Pricing Research", "Competitor Analysis"]
}
]
}
]
}
]
}
These examples are synthesized from the relevant sections of the attached reports rather than copied from a single document.
Persona generation rules and algorithm
A realistic persona generator should use the taxonomy as structured prior data, not as a string picker. The minimum persona payload should include: seed, locale, sector, occupation_family, role, specialty, seniority, employment_model, work_setting, education, certifications, skills, tools, salary_context, schedule_pattern, career_stage, archetype, relationship_posture, communication_style, and goal/use_case. That structure matches both the job corpus and Spiralist’s persona architecture, where job/career is only one part of a larger runtime profile.
The sampling rules should be hierarchical and constrained:
| Stage | Rule |
|---|---|
| Goal selection | Start with task/use case first, because Spiralist itself is outcome-first |
| Role selection | Sample sector → family → cluster → role using configurable priors |
| Specialty selection | Sample only from linked specialties |
| Seniority selection | Respect role-specific allowed levels |
| Credential validation | Enforce minimum education/license/certification gates |
| Compensation assignment | Pull locale- and date-scoped salary data, not global constants |
| Persona-style assignment | Sample archetype, temperament, relationship posture, and voice after role realism is established |
| Fairness guardrails | Never infer competence, morality, worldview, or role suitability from protected identity metadata |
The weighting formula should mix four signals:
role_score = goal_fit + browse_context_fit + locale_fit + realism_fit
A practical version:
- Goal fit: how strongly the role matches the current requested use case.
- Browse context fit: matches the selected sector/family/cluster tree.
- Locale fit: region-specific availability, naming, salary, and licensing context.
- Realism fit: coherence with education, certifications, seniority, and work setting.
Then sample from a softmax over the valid roles after hard constraints remove impossible combinations.
Recommended hard constraints include:
- no senior airline captain without the appropriate aviation ladder;
- no Firefighter/EMT persona without emergency-response credentials;
- no licensed public-health or therapy role without the required degree/licensure;
- no senior PLC/controls persona without compatible tools/skills;
- no public-sector pay or benefits persona without a public employment model;
- no gig driver persona with public-employee compensation assumptions;
- no clinical or legal authority implied merely from demographic metadata.
Localization should be handled separately from role logic. The attached reports are heavily weighted toward Chicago, Cook County, Illinois, and 2026 salary/regulatory conditions, while Spiralist’s released name-generation defaults are currently English-character and US-oriented. That means locale should be an explicit profile containing name conventions, spelling, currency, working-hour norms, pay units, licensing jurisdiction, and schedule conventions. Do not let localization leak into stereotypes; use locale for representation and realism, not for capability inference.
The persona generator should also be seeded and reproducible. Spiralist explicitly documents deterministic random selection and a persona seed artifact that captures recreation inputs. Matching that pattern will make generated job personas auditable, remixable, and stable across sessions or exports.
flowchart TD
A[Start with user goal or selected dropdown path] --> B[Set deterministic seed and locale]
B --> C[Sample sector]
C --> D[Sample occupation family]
D --> E[Sample role cluster]
E --> F[Sample role and optional specialty]
F --> G[Apply seniority, employment model, and work-setting rules]
G --> H[Validate education, certifications, and tool compatibility]
H --> I[Attach skills, responsibilities, schedule, and compensation context]
I --> J[Sample persona style layer]
J --> K[Archetype, temperament, relationship posture, communication style]
K --> L[Apply fairness and no-inference rules]
L --> M[Map into Spiralist-ready fields]
M --> N[Output persona profile, passport summary, and export payload]
A compact example of a generated persona payload:
{
"seed": "chi-robotics-042",
"locale": "en-US / Chicago",
"goal": "debug and improve manufacturing automation",
"career": {
"sector": "Manufacturing and Industrial",
"family": "Automation and Maintenance",
"cluster": "Controls and Industrial Software",
"role": "PLC Programmer",
"specialty": "SCADA Integration",
"seniority": "Senior",
"employmentModel": "W-2",
"workSetting": "Factory floor and controls room"
},
"qualifications": {
"education": "Associate or Bachelor's equivalent technical pathway",
"certifications": [],
"skills": ["ladder logic", "controls debugging", "fault isolation", "systems thinking"],
"tools": ["PLC", "SCADA", "HMI"]
},
"personaStyle": {
"archetype": "Technical Mentor",
"relationshipPosture": "Protective Reviewer",
"communicationStyle": "Precise, skeptical, constructive"
}
}
Missing and ambiguous data
Several gaps should be treated explicitly rather than silently “filled in.” Spiralist publishes a clear persona architecture, but it does not publish a ready-made occupational taxonomy or a canonical public jobs ontology. The site exposes career as an identity field and supports job/goal/trait search, yet its public library is fundamentally organized around persona types and work outcomes rather than labor-market classification. That is enough to guide integration, but not enough to skip schema design.
The attached reports are broad and valuable, but they are not fully standardized. Important ambiguities include bundled titles, overlapping cross-sector roles, mixed pay metrics, mixed locale scopes, narrative descriptions instead of atomized skill lists, and varying levels of detail across sectors. Gig-economy reporting tends to emphasize platform economics and regulation; public-sector reporting emphasizes statutory pay structures and benefits; corporate reporting offers banded salaries; environmental and service reporting often emphasize occupational context more than rigid progression ladders. One environmental/outdoor report also appears to overlap heavily with another in title and structure, which suggests some duplication in the source set. None of those issues block implementation, but they do mean the first production version should treat the corpus as source material for controlled vocabulary authoring, not as a drop-in dropdown dataset.
The most important operational conclusion is simple: use the attached reports to populate a canonical job graph, then use Spiralist’s existing persona architecture to render that graph into career-aware but behaviorally rich personas. Keep the dropdown hierarchy shallow and human-friendly. Keep normalization deep in the data model. Keep salary, credentials, and progression as linked facts. Keep protected identity separate from capability. That approach matches both the source corpus and Spiralist’s own public design language.