AI4Portfolio: A Design-Science-Informed, Human-in-the-Loop Reference Architecture for AI-Assisted Portfolio Management in Refining
Danielle Fernandes prestígio (8)
Avaliação por modelos + Supervisão Editorial.
Aceite provisório
Registro já válido e público — número, score e prova congelados. Em revisão complementar do comitê; um endosso o torna definitivo, e o número não muda.
Arquitetura com IA generativa promete organizar portfólios industriais sem tirar a decisão das mãos humanas
Resumo executivo
O artigo apresenta o AI4Portfolio, uma arquitetura de referência orientada por design science para inserir IA generativa no front end da gestão de portfólio industrial em refino, mantendo rastreabilidade e autoridade decisória humana. A solução integra registros estruturados no Azure DevOps, gates ponderados técnico-financeiro-estratégicos, geração de business cases a partir de transcrições, recomendações de veículos de inovação e financiamento, e governança human-in-the-loop.
A demonstração retrospectiva reprocessou 26 projetos do portfólio de 2025 e comparou estimativas internas agregadas de esforço em três atividades de preparação de business case. O próprio texto delimita que se trata de evidência formativa de viabilidade operacional, não de teste causal de efetividade. A principal contribuição declarada é o pacote transferível: modelo em seis camadas, lógica explícita de pontuação, esquema funcional de prompts e saídas, guardrails de validação humana e protocolo prospectivo de avaliação.
Paper completo
Entre para baixar o PDF registradoAI4Portfolio: A Design-Science-Informed, Human-in-the-Loop Reference Architecture for AI-Assisted Portfolio Management in Refining Abstract AI4Portfolio is presented as a design-science-informed reference architecture for embedding generative artificial intelligence (GenAI) in the front end of industrial portfolio management while preserving human decision authority.
The artifact was implemented at Acelen in a refining context and integrates five elements: structured portfolio records in Azure DevOps, weighted technical/financial/strategic decision gates, transcript-to-business-case generation, recommendations of innovation vehicles and funding mechanisms, and human-in-the-loop governance.
The study is deliberately framed as an artifact design and formative demonstration, not as a causal effectiveness trial. A retrospective demonstration reprocessed 26 projects from the 2025 portfolio and compared rounded, aggregated internal effort estimates for three business-case activities. The reported estimates changed from 2 to 1 h/project for problem framing, 4 to 2 h/project for assessment of comparable market solutions, and 8 to 2 h/project for return structuring. These differences correspond to descriptive reductions of 50%, 50%, and 75%, respectively, and a combined change from 14 to 5 h/project (64.3%). Because project-level timestamps, randomization, a control group, blinding, and direct instrumented time measurement were not available, these numbers are treated only as formative operational estimates and cannot establish causality. The scientific contribution is instead the transferable architecture and replication package: a platform-agnostic six-layer model, explicit scoring logic, a functional prompt schema, human-validation guardrails, and a prospective evaluation protocol. This reframing separates what was implemented and demonstrated from what remains to be validated, and provides a reusable structure for organizations seeking governed GenAI support for portfolio decision processes.
Keywords: design science; strategic portfolio management; project portfolio management; generative AI; humanin-the-loop; decision support; refining; innovation governance.
1. Introduction
Digital transformation in asset-intensive industries depends on the ability to convert a high volume of heterogeneous technological demands into a governed portfolio of comparable initiatives. In refining, this front end is particularly demanding because proposals often combine process, digital, cybersecurity, operational, financial, and innovation dimensions. The management problem is therefore not only to generate ideas, but also to structure evidence, compare alternatives, assess feasibility and value, and move initiatives through decision gates with sufficient traceability.
Project portfolio management (PPM) research treats portfolio selection as a recurrent, multi-stage decision problem involving strategic alignment, value, risk, resource constraints, and interdependencies. Archer and Ghasemzadeh (1999) proposed an integrated staged framework for portfolio selection; Cooper, Edgett and Kleinschmidt (2001) showed the value of combining financial and non-financial criteria; Meskendahl (2010) linked portfolio configuration with strategy implementation; and Martinsuo (2013) emphasized that portfolio decisions are context-dependent rather than purely optimization problems. Multi-criteria decision analysis (MCDA) research similarly supports the explicit combination of multiple decision attributes instead of reliance on a single return metric (Danesh, Ryan and Abbasi, 2018; Kandakoglu, Walther and Ben Amor, 2024).
Generative AI adds a new capability to this setting: unstructured evidence such as meeting transcripts, narratives, and technical descriptions can be transformed into structured governance artifacts. Yet the existence of a language model does not by itself create a scientifically meaningful portfolio-management contribution. Müller et al. (2024) argue that AI-related research in project management must clearly separate technological novelty from sound empirical evaluation. Human-AI research also shows that oversight must be deliberately designed because human reviewers may accept incorrect automated recommendations or over-trust generated outputs (Westphal et al., 2023).
This paper therefore reframes AI4Portfolio as a design artifact rather than as a claim of algorithmic novelty or proven causal productivity gain. Design Science Research (DSR) is appropriate for this purpose because it focuses on creating and evaluating artifacts intended to improve human and organizational capabilities (Hevner et al., 2004). The Design Science Research Methodology (DSRM) proposed by Peffers et al. (2007) further structures this process into problem identification, objective definition, design and development, demonstration, evaluation, and communication.
The research question is: How can a generative-AI layer be designed and embedded in an industrial portfoliogovernance process so that unstructured project information is converted into standardized, traceable decision artifacts without transferring decision authority from humans to the AI system? The paper contributes (i) a platformagnostic reference architecture; (ii) explicit decision and traceability rules; (iii) a reproducible functional prompt and output schema; and (iv) a formative industrial demonstration that is reported with explicit limits on causal interpretation.
2. Theoretical Background and Positioning
2.1 Portfolio governance as a staged and multi-criteria process Portfolio selection is not a single scoring event. It involves successive information transformations and decisions, from opportunity identification to feasibility, prioritization, authorization, and execution. This supports an architecture in which each stage has a defined purpose, required evidence, scoring rule, and accountable decision body.
Weighted scoring is used in AI4Portfolio as a transparent aggregation mechanism, not as an assertion that weights are universally optimal. The literature on MCDA and portfolio selection supports explicit criteria and aggregation while also warning that weights and scales must remain visible to decision makers and be interpreted within organizational context.
2.2 AI-assisted portfolio decision support and human control GenAI can reduce the manual burden of converting unstructured project information into standardized artifacts, but it also introduces risks of hallucination, unsupported benchmarking, inconsistent financial assumptions, prompt sensitivity, and automation bias. Consequently, the AI component is positioned as an advisory transformation layer rather than an autonomous decision maker.
In AI4Portfolio, the assistants may extract, organize, classify, summarize, and recommend. They may not approve investments, alter gate weights autonomously, or convert an unsupported assumption into an accepted fact.
Decision authority remains with technical, management, and executive committees.
2.3 Positioning relative to commercial SPM platforms Commercial Strategic Portfolio Management (SPM) platforms already offer AI-supported reporting, prioritization, scenario analysis, resource/risk insights, and conversational assistance. AI4Portfolio is therefore not positioned as a competing general-purpose SPM product. Its specific contribution is an internal orchestration pattern that links transcript-derived evidence, domain-specific gates, business-case structuring, innovation/funding recommendations, a pre-existing enterprise system of record, and explicit human approval.
Dimension Commercial SPM platforms AI4Portfolio Research implication Scope Enterprise strategy-toexecution and portfolio management Front-end structuring of innovation/digital initiatives in refining Domain-specific integration rather than new SPM category AI use Insights, reporting, forecasting, prioritization, workflows Transcript-to-artifact structuring;
benchmarking/return support;
innovation/funding guidance Specific transformation of unstructured evidence into gate-ready artifacts System of record Native platform data model Azure DevOps work items and dashboards Architecture can be decoupled from the chosen software Decision authority Implementationdependent Explicit technical, management and executive gates Human-in-the-loop is a design rule, not an optional interface feature Novelty claim Product capability Governance artifact and reference architecture Contribution is integrative and incremental, not algorithmic
3. Methodology
3.1 Design-science-informed research approach The study adopts a design-science-informed approach. The primary object of evaluation is the utility and coherence of the AI4Portfolio artifact within a real portfolio-governance process. The 26-project reprocessing exercise is treated as a formative demonstration of operational feasibility, not as a controlled test of causal productivity effects.
DSRM activity AI4Portfolio realization Evidence in this paper
1. Problem identification
Long and heterogeneous front-end structuring activities; fragmented evidence; repeated business-case preparation Industrial problem description and IPM workflow
2. Objectives of a solution
Standardize information, preserve traceability, reduce manual structuring burden, keep human decision authority Explicit design objectives and governance constraints
3. Design and development
Weighted gates + structured work items + two GenAI assistants + dashboards Architecture, scoring equation, input/output schemas and guardrails
4. Demonstration
Reprocessing of 26 historical portfolio projects Descriptive effort estimates and example artifacts
5. Evaluation
Formative, retrospective, aggregated comparison No causal inference; limitations and prospective protocol are explicit
6. Communication
Documentation of transferable architecture and replication package Sections 3.8 and Appendices A-B 3.2 Integrated Portfolio Management (IPM) governance flow The IPM framework connects an Upstream phase—idea capture, synthesis, technical evaluation, management analysis, and executive decision—with a Downstream phase for design, execution, and feedback. Each stage has a defined decision purpose, required information, and accountable governance body.
Figure 1. Integrated Portfolio Management macro-flow: Upstream maturation and Downstream execution.
Stage 1 – Idea Capture. Identify business opportunities and challenges. Minimum information includes the problem, proposed solution, expected benefits, business unit, time horizon, and sponsor.
Stage 2 – Synthesis (70%). Assess strategic alignment, sponsorship, data availability, business-process impact, business value, productivity, non-financial value, and risk mitigation.
Stage 3 – Technical Committee / Technical Gate (10%). Assess architecture alignment, integration complexity, solution flexibility, internal/external capability, dependencies, risks, and Technology Readiness Level (TRL).
Stage 4 – Management Analysis and Committee (20%). Assess prioritization using financial impact, risk mitigation, return, payback, strategic synergy, sustainability, and consolidated indicators.
Stage 5 – Executive Committee. Approve or reject initiatives requiring executive authorization based on the consolidated business case, financial information, KPIs, portfolio position, and execution roadmap.
3.3 Scoring model and traceability Final prioritization combines the weighted stage scores from Synthesis (70%), Technical Gate (10%), and Management Committee (20%). If S_S, S_T, and S_M denote the stage-level scores on the 0–3 scale, the final score is:
푺 풇풊풏풂풍 = ퟎ.ퟕퟎ·푺 푺 + ퟎ.ퟏퟎ·푺 푻 + ퟎ.ퟐퟎ·푺 푴
For communication, the result can be normalized to 0–100 as 100·S_final/3. For example, stage scores 2.4, 2.0, and 2.5 yield S_final = 2.38, equivalent to 79.3/100. This numerical example is illustrative and is not one of the 26 demonstration projects.
Score Operational anchor Interpretation Absent, incompatible, or unsupported by evidence Mandatory information is missing or a material constraint blocks advancement; flag is mandatory.
1 Weak / incomplete Evidence exists but substantial gaps, uncertainty, or corrective work remain.
2 Adequate / partially validated Evidence is sufficient for continued analysis, with manageable assumptions or gaps requiring validation.
3 Strong / decision-ready Evidence is traceable and consistent with the criterion; no material gap is identified at the current gate.
The score is therefore not produced by the LLM alone. Each criterion must be supported by structured evidence in the portfolio record. Missing required information receives the minimum score and is flagged for completion.
Committees may override an AI-supported preliminary classification, but the final decision and rationale remain attributable to human governance.
3.4 Prioritization matrix and digital traceability Initiatives are registered in Azure DevOps as work items categorized as projects. The work item acts as the auditable record for problem definition, solution, sponsor, strategic drivers, technical attributes, financial metrics, benefits, assumptions, and decision status. Structured records feed dashboards and a value-versus-viability prioritization matrix. The upper-right region represents high-value/high-viability initiatives (Results Acceleration / Strategic Opportunities); the upper-left region represents high-value/lower-viability Strategic Bets / Disruptive Innovation; the lower region represents Incremental or Marginal Gains. This description is included so that the method remains understandable even when confidential figures are blurred.
Figure 2a. Example of a structured Azure DevOps work item used as the portfolio data source.
Figure 2b. Portfolio prioritization matrix showing value versus viability.
Figure 2c. Automatically generated business-case dashboard. Image clarity is intentionally reduced to preserve confidential information.
3.5 AI4Portfolio assistants and reproducibility boundary Two AI assistants are integrated into Acelen GPT. Reproducibility is documented at two levels: (i) process-level reproducibility, which is provided in this paper through the functional input/output and governance specification; and (ii) output-level reproducibility, which would require the exact model build, sampling parameters, and verbatim proprietary prompts.
Assistant Required input Structured output Mandatory validation Problem and Solution Meeting transcript and available project context Problem statement;
proposed solution;
tangible/intangible benefits; preliminary gate inputs; comparable solutions;
investment/return hypotheses Unsupported content must be flagged;
benchmarks and financial assumptions require human verification.
Innovation Projects Project characteristics, maturity, constraints and expected return Recommended innovation vehicle and potential funding mechanisms Recommendation is reviewed by portfolio specialists; final vehicle/funding decision remains human.
The research dataset available for this manuscript does not contain a stable record of the exact backend model/version, temperature (or equivalent sampling settings), or the verbatim internal system prompts used during the 2025 retrospective demonstration. It would therefore be misleading to claim deterministic output-level reproducibility. To reduce this limitation, Appendix A provides vendor-neutral canonical prompts and an output schema that reproduce the intended task logic and guardrails without disclosing proprietary internal instructions.
Replication of the governance method is therefore possible; replication of identical historical LLM outputs is not.
Figure 3a. Example output from the Problem and Solution assistant.
Figure 3b. Example output from the Innovation Projects assistant.
3.6 Retrospective formative demonstration The demonstration sample comprises 26 projects from the 2025 portfolio. The same historical projects were reprocessed with the assistants, and the resulting artifacts were compared operationally with the business-case process previously used for those initiatives.
The retained research dataset contains only rounded, aggregated internal effort estimates for three activities:
problem understanding and framing; assessment of comparable market solutions; and structuring of expected return and financial justification. It does not contain project-level time stamps, individual before/after observations, or an auditable contemporaneous time-measurement protocol. The values are therefore treated in this paper as retrospective operational estimates, not as instrumented measurements.
The same projects had already been processed by the organization before AI-assisted reprocessing. Analyst familiarity with the cases is therefore a material confounding factor. No control group, randomization, blinding, or project-complexity stratification was used. The demonstration answers whether the artifact can be used in the workflow and whether the reported effort estimates are directionally lower; it does not identify the causal effect of AI.
The descriptive percentage difference is calculated as:
푫풆풔풄풓풊풑풕풊풗풆 풅풊풇풇풆풓풆풏풄풆 (%) = ( 푬풔풕풊풎풂풕풆 풃풆풇풐풓풆 − 푬풔풕풊풎풂풕풆 풂풇풕풆풓 ) 푬풔풕풊풎풂풕풆 풃풆풇풐풓풆 × ퟏퟎퟎ Because the underlying project-level observations are unavailable, no standard deviation, confidence interval, hypothesis test, or effect-size statistic is reported. Reporting such statistics from the three rounded averages would create false precision.
3.7 Prospective effectiveness-evaluation protocol A future effectiveness study should be prospectively designed and should separate speed from artifact quality. New initiatives should be randomized or, where randomization is operationally infeasible, matched by complexity and business area. Workflow logs should capture start/finish timestamps automatically; reviewers should be blinded to whether an artifact was AI-assisted; and at least two qualified reviewers should independently score artifact quality.
Quality criterion Observable evidence 0–3 anchor Completeness Presence of required gate and business-case fields 0 absent; 1 major gaps; 2 minor gaps; 3 complete Traceability to source Whether statements are supported by transcript or registered project evidence 0 material contradiction; 1 several unsupported claims; 2 mostly traceable; 3 fully traceable Financial coherence Consistency among assumptions, investment, benefit, payback/return logic 0 material error; 1 major correction;
2 minor correction; 3 coherent Benchmark verifiability Ability to verify comparable solutions and cited sources 0 unverifiable; 1 weak; 2 mostly verifiable; 3 fully verifiable Strategic/technical consistency Alignment between proposed solution and governance criteria 0 inconsistent; 1 weak; 2 adequate;
3 strong Decision readiness Expert rework required before committee use 0 rebuild; 1 substantial rework; 2 minor edits; 3 ready
Additional downstream outcomes should include rework rate, approval cycle time, forecast accuracy, realized value, human override frequency, and incidence of factual/financial corrections. This protocol is proposed for future validation and was not applied retrospectively to the 2025 sample.
3.8 Transferable reference architecture To make the contribution transferable beyond Acelen, AI4Portfolio is abstracted into six layers. None of the layers requires Azure DevOps, Acelen GPT, or the specific company governance nomenclature. The local implementation is one instantiation of the reference architecture.
Layer Acelen implementation Transferable abstraction Minimum replication requirement
1. Evidence capture
Meeting transcripts and project inputs Unstructured evidence intake Source material with ownership, date and project identifier
2. AI transformation Acelen GPT assistants
Controlled LLM transformation service Versioned task instructions; structured output schema;
refusal/uncertainty rules
3. System of record Azure DevOps work item Traceable portfolio record
Unique project ID;
required fields; version history; source links
4. Decision model
70/10/20 weighted stages and 0–3 criteria Explicit multi-criteria gate model Published criteria, scales, weights, missing-data rule and aggregation equation
5. Human governance
Technical, management and executive committees Accountable human decision authority Named decision roles;
override rights; recorded rationale
6. Feedback and
assurance Dashboards, monitoring and updates Closed-loop portfolio learning Execution feedback, audit trail, error/rework logging and periodic model review
The architecture can therefore be instantiated with alternative tools: a document-management platform may replace Azure DevOps; a different enterprise LLM may replace Acelen GPT; and gate weights may be redesigned for another sector. What must remain invariant for conceptual replication is the separation of evidence, AI transformation, traceable records, explicit decision logic, accountable human authority, and feedback.
4. Formative Demonstration Results
4.1 Descriptive effort estimates Table 1 reports only the rounded aggregated estimates retained from the 26-project retrospective exercise. The terminology “estimate” is used intentionally because the dataset does not contain project-level time logs or a validated measurement protocol.
Activity Before estimate (h/project) AI-assisted estimate (h/project) Difference (h/project) Descriptive difference Indicative aggregate difference (N=26) Problem understanding and framing 2 1 1 50.0% 26 h Comparable market-solution assessment 4 2 2 50.0% 52 h Return structuring and financial justification 8 2 6 75.0% 156 h Total across the three activities 14 5 9 64.3% 234 h
The return-structuring difference is 75.0% because the reported estimate changes from 8 to 2 h/project. The earlier value of 62.5% was arithmetically inconsistent with the reported hours. Across all three activities, the rounded estimates sum to 14 h/project before and 5 h/project in the AI-assisted reprocessing, a descriptive difference of 9 h/project (64.3%). Multiplying this difference by 26 yields 234 indicative hours; this extrapolation is arithmetic, not a statistical estimate of population savings.
4.2 What the demonstration does and does not show The demonstration shows that the AI4Portfolio artifact can be operationally integrated into the existing portfolio workflow and can generate the expected structured outputs for historical cases.
The direction of the rounded effort estimates is consistent with the intended automation target: the largest difference appears in return structuring, the activity with the highest reported manual effort.
The demonstration does not show that AI caused the difference, that business-case quality improved, or that portfolio decisions became financially better. Familiarity with the same projects, standardization effects, analyst learning, and unmeasured case complexity are plausible alternative explanations.
5. Discussion
5.1 Nature of the contribution The contribution of AI4Portfolio is incremental but explicit: it is a reusable governance and information architecture for controlled GenAI use in portfolio front-end work. The artifact does not introduce a new optimization algorithm, language model, or MCDA method. Its contribution lies in how established components are composed into an auditable human-AI decision process and how that composition is specified for replication.
This is consistent with the logic of design science, in which novelty may reside in an artifact that solves a relevant organizational problem through a novel or useful configuration of technological and organizational elements, provided that the artifact and its utility are clearly described (Hevner et al., 2004; Peffers et al., 2007).
5.2 Transferability beyond Acelen The six-layer reference architecture in Section 3.8 is the main transfer mechanism. It distinguishes companyspecific instantiation from invariant design principles. Another organization can replicate the method without using Acelen’s software stack by mapping its own tools and governance roles to the six layers. This makes the contribution more general than a case-specific workflow, while avoiding the stronger and unsupported claim that the numerical effort estimates will generalize across organizations.
Design principle Why it is transferable Failure mode if omitted Evidence before recommendation Any portfolio can preserve provenance between source and generated artifact Hallucinated or unsupported claims become indistinguishable from evidence Structured output contract Any LLM can be required to populate a known schema Generated text becomes difficult to compare, audit, or score Explicit gate mathematics Any organization can publish its own scales and weights Prioritization becomes opaque or model-dependent Human decision authority Applies independently of sector or software Automation bias can convert advisory output into unaccountable approval Versioning and feedback Applies to any evolving AI system Model/prompt changes cannot be linked to changes in output quality 5.3 Reproducibility: conceptual versus output-level The revised manuscript distinguishes two reproducibility targets. Conceptual/process reproducibility asks whether another organization can recreate the workflow, schemas, scoring logic, guardrails, and evaluation protocol; the manuscript now supports this through Sections 3.3–3.8 and Appendices A-B. Output-level reproducibility asks whether the same transcript will generate the same text; this cannot be established for the historical demonstration because the exact model build, proprietary prompt text, and sampling configuration were not retained in the research dataset. The distinction is important because claiming the second form of reproducibility would overstate the evidence.
5.4 Methodological robustness and evidence maturity The methodological limitation cannot be repaired retrospectively by adding statistical language to aggregated estimates. The defensible strategy is to align the claim with the evidence. Accordingly, this paper treats the 26- project exercise as formative demonstration evidence (artifact feasibility and directional operational signal) rather than summative effectiveness evidence. Causal effectiveness, quality equivalence, and generalizable productivity gains remain untested hypotheses for the prospective protocol in Section 3.7.
5.5 AI risks and governance Generative-AI outputs may contain fabricated benchmarks, incorrect cost or benefit assumptions, inconsistent classifications, or confident statements unsupported by the transcript. The reference architecture therefore requires source traceability, explicit uncertainty, human validation of financial and benchmarking content, version control, and logging of overrides. Human-in-the-loop is treated as an accountability mechanism, not simply as a final confirmation button.
5.6 Limitations
• The 26-project demonstration is retrospective and reuses projects previously known to the organization,
creating a strong familiarity/learning confound.
• The research dataset contains rounded aggregated effort estimates rather than project-level instrumented
observations; the original time-measurement provenance cannot be independently reconstructed from the data available for this manuscript.
- No contemporaneous control group, randomization, blinding, or complexity stratification was used.
- Artifact quality was not evaluated with a blinded rubric in the retrospective demonstration.
- The exact historical model/version, sampling settings, and verbatim proprietary prompts were not retained in
the research dataset, preventing deterministic output-level replication.
• The study does not measure realized portfolio value, ROI accuracy, approval quality, or downstream execution
outcomes.
• The transferable architecture is analytically generalized from one organizational implementation; crosscompany empirical validation remains future work.
6. Conclusion
AI4Portfolio is a design-science-informed reference architecture for embedding GenAI in the front end of industrial portfolio management while preserving traceability and human decision authority. Its six layers separate evidence capture, AI transformation, system of record, decision logic, human governance, and feedback/assurance.
The retrospective 26-project exercise provides only formative evidence. Rounded internal estimates are directionally lower in the AI-assisted reprocessing—50% for problem framing, 50% for comparable-solution assessment, and 75% for return structuring—but the design does not permit causal attribution. These numbers should therefore be read as operational signals motivating a prospective evaluation, not as statistically proven productivity effects.
The main contribution is the transferable design package: explicit gate mathematics, structured AI input/output contracts, source-traceability rules, human-validation guardrails, a platform-agnostic reference architecture, and a prospective effectiveness protocol. Future work should validate the architecture on new projects with direct workflow timestamps, complexity controls, blinded artifact-quality scoring, model/prompt versioning, and downstream business outcomes.
7. Acknowledgments
The authors acknowledge the Acelen Technology team for collaboration in defining technical evaluation criteria and for providing the digital assistants used in business-case generation and prioritization. The authors also thank the Acelen business-unit teams for supporting the execution and continuous improvement of the portfolio process.
Appendix A. Vendor-Neutral Replication Prompts and Output Contract The following templates are not verbatim Acelen proprietary prompts. They are canonical, vendor-neutral prompts derived from the implemented task logic and are provided so that the method can be conceptually reproduced on another enterprise LLM.
A.1 Problem and Solution assistant – canonical prompt ROLE: You are an industrial portfolio analyst supporting decision preparation, not an investment approver.
INPUTS: (1) meeting transcript; (2) project identifier; (3) known business context; (4) available technical/financial data; (5) governance criteria.
RULES: Use only supplied evidence for factual statements. Separate evidence from inference. Never invent benchmark companies, costs, revenues, TRL, or financial results. If evidence is missing, write “verification required”. For every material statement, identify the source field or transcript segment. Financial estimates must list assumptions and formulas.
TASKS: 1) formulate the problem; 2) formulate the proposed solution; 3) classify tangible and intangible benefits; 4) identify missing information; 5) suggest preliminary gate inputs without making the final decision; 6) list comparable-solution hypotheses requiring verification; 7) structure investment/return assumptions.
OUTPUT: return the structured fields defined in Table A1 and a final “human validation required” section.
A.2 Innovation Projects assistant – canonical prompt ROLE: You are an innovation-portfolio advisor.
INPUTS: project problem, proposed solution, TRL/maturity, required capabilities, expected value, time horizon, intellectual-property needs, budget constraints, and partnership constraints.
RULES: Do not invent public funding eligibility. Distinguish confirmed requirements from hypotheses.
Provide rationale and required verification for every recommendation.
TASKS: rank plausible innovation vehicles (internal development, vendor, startup, ICT/R&D collaboration, co-development, other) and identify candidate funding mechanisms only when eligibility can be justified from supplied information.
OUTPUT: ranked recommendations, rationale, assumptions, missing information, validation actions, and confidence/uncertainty statement.
A.3 Minimum output schema Field Type Requirement project_id string Mandatory unique identifier problem_statement text Must be traceable to source evidence proposed_solution text Must distinguish proposal from observed fact benefits_tangible list Each benefit labeled evidence/estimate/assumption benefits_intangible list Same provenance rule missing_information list Mandatory when any required field is unsupported benchmark_hypotheses list Must be marked verification required unless a verifiable source is attached financial_assumptions table Assumption, value, source, owner, validation status preliminary_gate_inputs structured values Advisory only; no autonomous approval source_trace list Reference to transcript segment or registered field human_validation_required boolean + notes Always true for benchmark, financial and final prioritization outputs Appendix B. Replication and Evaluation Checklist
- Define the local portfolio stages, decision rights, criteria, weights, and 0–3 (or equivalent) operational anchors.
- Define a system-of-record schema with unique project IDs, required fields, source links, and version history.
- Implement the canonical prompt logic or an equivalent controlled prompt and require a structured output
contract.
4. Version the LLM model identifier, prompt, retrieval configuration, temperature/sampling parameters, and
deployment date for every evaluation cycle.
5. Preserve source-to-output traceability and require explicit “verification required” markers for unsupported
information.
- Pilot on new—not previously processed—projects and capture workflow start/finish timestamps automatically.
- Stratify or match projects by complexity, business area, and artifact type; use a contemporaneous non-AI
comparison when feasible.
- Use at least two blinded reviewers to score artifact quality and report inter-rater agreement.
- Report project-level distributions (mean, median, standard deviation, range, confidence interval) rather than
only rounded averages.
- Report error, hallucination, financial-correction, human-override, and rework rates.
- Measure downstream outcomes such as approval cycle time, forecast accuracy, realized value, and decision
reversals.
12. Document which elements are organization-specific and which correspond to the six invariant layers of the
reference architecture.
8. References
Archer, N. P., & Ghasemzadeh, F. (1999). An integrated framework for project portfolio selection. International Journal of Project Management, 17(4), 207–216. https://doi.org/10.1016/S0263-7863(98)00032-5 Artto, K., Kulvik, I., Poskela, J., & Turkulainen, V. (2011). The integrative role of the project management office in the front end of innovation. International Journal of Project Management, 29(4), 408–421.
https://doi.org/10.1016/j.ijproman.2011.01.008 Cooper, R., Edgett, S., & Kleinschmidt, E. (2001). Portfolio management for new product development: Results of an industry practices study. R&D Management, 31(4), 361–380. https://doi.org/10.1111/1467-9310.00225 Danesh, D., Ryan, M. J., & Abbasi, A. (2018). Multi-criteria decision-making methods for project portfolio management: A literature review. International Journal of Management and Decision Making, 17(1), 75–94.
https://doi.org/10.1504/IJMDM.2018.088813 Felicetti, A. M., Cimino, A., Mazzoleni, A., & Ammirato, S. (2024). Artificial intelligence and project management: An empirical investigation on the appropriation of generative chatbots by project managers. Journal of Innovation & Knowledge, 9(3), 100545. https://doi.org/10.1016/j.jik.2024.100545 Hevner, A. R., March, S. T., Park, J., & Ram, S. (2004). Design science in information systems research. MIS Quarterly, 28(1), 75–105. https://doi.org/10.2307/25148625 Kandakoglu, M., Walther, G., & Ben Amor, S. (2024). The use of multi-criteria decision-making methods in project portfolio selection: A literature review and future research directions. Annals of Operations Research, 332, 807–
830. https://doi.org/10.1007/s10479-023-05564-3
Martinsuo, M. (2013). Project portfolio management in practice and in context. International Journal of Project Management, 31(6), 794–803. https://doi.org/10.1016/j.ijproman.2012.10.013 Meskendahl, S. (2010). The influence of business strategy on project portfolio management and its success—A conceptual framework. International Journal of Project Management, 28(8), 807–817.
https://doi.org/10.1016/j.ijproman.2010.06.007 Müller, R., Locatelli, G., Holzmann, V., Nilsson, M., & Sagay, T. (2024). Artificial Intelligence and Project Management: Empirical Overview, State of the Art, and Guidelines for Future Research. Project Management Journal, 55(1), 9–15. https://doi.org/10.1177/87569728231225198 Peffers, K., Tuunanen, T., Rothenberger, M. A., & Chatterjee, S. (2007). A Design Science Research Methodology for Information Systems Research. Journal of Management Information Systems, 24(3), 45–77.
https://doi.org/10.2753/MIS0742-1222240302 Taheripour, E., & Sadjadi, S. J. (2026). Project portfolio management in the age of artificial intelligence: A review of challenges, key features, and future research directions. Journal of Project Management, 11(1), 247–272.
https://doi.org/10.5267/j.jpm.2025.9.006 Westphal, M., Vössing, M., Satzger, G., Yom-Tov, G. B., & Rafaeli, A. (2023). Decision control and explanations in human-AI collaboration: Improving user perceptions and compliance. Computers in Human Behavior, 144,
107714. https://doi.org/10.1016/j.chb.2023.107714
Como citar
FERNANDES, Danielle. AI4Portfolio: A Design-Science-Informed, Human-in-the-Loop Reference Architecture for AI-Assisted Portfolio Management in Refining. Crivum, 2026. Disponível em: https://crivum.org/p/qed8pfg0.
Fernandes, D. (2026). AI4Portfolio: A Design-Science-Informed, Human-in-the-Loop Reference Architecture for AI-Assisted Portfolio Management in Refining. Crivum. https://crivum.org/p/qed8pfg0
@misc{crivum_ma77,
author = {Danielle Fernandes},
title = {AI4Portfolio: A Design-Science-Informed, Human-in-the-Loop Reference Architecture for AI-Assisted Portfolio Management in Refining},
year = {2026},
publisher = {Crivum},
url = {https://crivum.org/p/qed8pfg0},
note = {Crivum MA77},
}Viu um problema nesta obra?
Toda obra publicada aqui fica aberta ao exame de quem lê. Se você encontrou dado que não se sustenta, trecho sem crédito ou autoria que não confere, aponte. A editoria lê todos os apontamentos.
Apontar exige conta, para o registro ter um nome por trás. Entre antes de escrever, assim você não perde o texto. Entrar para apontar
Um apontamento vai direto à editoria e fica registrado. Quando 3 pessoas diferentes apontam a mesma obra, ela sai do ar na hora e o caso vai ao comitê. A avaliação e o registro permanecem em qualquer caso, o que muda é a disponibilidade.