Extend an Existing Evidence-Based Organization Profile & Capability Extraction System

USD 10–30

OpenListed onFreelancer.com
Fixed

About the project

We are looking for an experienced Python / document intelligence / data modeling engineer to extend an existing evidence-based organization profiling system. This is not a greenfield project. We already have a working Python codebase with: an existing typed Ngo / organization model; existing organization profile fields; evidence/provenance concepts; an evidence-backed capability projection layer; automated tests; anti-leakage safeguards; an existing KRS parser; an existing canonical organization schema; downstream decision logic already consuming organization data; one real structured organization capability already working end-to-end. Your task is to extend the existing implementation, not replace or redesign it. The main goal is to make the current organization profile useful enough for real decision-making by safely adding the most important missing organization capabilities from structured sources, statutes, questionnaires and other supplied documents. What already exists The current organization model already contains fields such as: organization type; registration date / year; statutory purposes; activity domains; target groups; geographic capacity; institution access; organization experience; staff capabilities; staff experience; program assets. The existing system also already includes: provenance-aware candidate concepts; UNKNOWN semantics; fail-closed behavior; deterministic tests; safeguards preventing unrelated free text from silently becoming a stable organization fact; downstream integration with an existing decision engine. You will receive the relevant source files, models, interfaces, tests and examples. You are not expected to design the organization model from scratch. Main Task Your task is to extend the current evidence-backed organization capability layer so that it can safely populate a bounded set of high-value organization profile fields. Possible priority fields include: statutory purposes; activity areas; target groups; geographic operating capacity; organization experience; institution / facility access; staff capabilities; staff experience or qualifications; program or operational assets; registration-related facts where authoritative evidence is available. The final v0.1 scope will be agreed before implementation. We do not expect support for every possible organization fact. Input Sources The module will work with an existing versioned intake contract: ORGANIZATION_INTAKE_V1 Inputs may include: organization name; Polish KRS registration number; structured questionnaire answers; statute / governing document; optional additional documents; structured facts already available in the existing system. The exact schema and existing code interfaces will be provided. Evidence Status One of the most important requirements is to keep different evidence levels clearly separated. The system should distinguish between: VERIFIED — supported by an authoritative structured source; DOCUMENT_EVIDENCED — explicitly supported by an organization document; DECLARED — supplied through a structured questionnaire; UNKNOWN — insufficient evidence. For example: Structured registry field → VERIFIED Statute, page 5 → DOCUMENT_EVIDENCED Questionnaire answer → DECLARED No reliable evidence → UNKNOWN A declaration must never silently become VERIFIED. Evidence-First Processing Conceptually: Structured source / statute / questionnaire / supplied document → evidence-backed candidate → validation → existing organization model Every positive fact should preserve appropriate provenance, such as: source type; document identifier/hash; page number where applicable; supporting evidence quote; source field/value; extractor or projector version; evidence status. Provenance should remain separate from stable organization semantics where appropriate. Important Anti-Leakage Rule The system must not derive stable organization capabilities from unrestricted unrelated free text. For example, information found only in: previous application text; project descriptions; marketing copy; notes; historical free-text profiles; unrelated documents; must not silently become a VERIFIED or stable capability. If a fact cannot be safely established, the correct result is: UNKNOWN not an inferred best guess. Document Processing Some supplied documents, particularly statutes and registration-related documents, will be in Polish. Native Polish fluency is helpful but not mandatory if you are comfortable working with Polish-language documents using modern LLMs, translation tools and provided acceptance cases. We will provide representative documents and expected outputs. LLM Usage We are open to LLM-assisted extraction where it is useful, especially for long statutes or complex documents. However: LLM output must be treated as a candidate, not automatically as truth; structured/schema-constrained output is preferred; evidence must be preserved; deterministic validation should be applied where possible; unsupported or ambiguous information must remain UNKNOWN; the implementation should not unnecessarily depend on a single LLM provider. A hybrid approach using deterministic parsing + LLM extraction + validation is welcome. Important Architectural Constraints Please do not redesign the existing system. In particular: do not replace the existing organization model; do not create a parallel profile architecture; do not redesign the database unless explicitly approved; do not build a frontend; do not build the intake form; do not implement n8n orchestration; do not modify the opportunity/document extraction module; do not rewrite the downstream decision engine; do not perform unrelated refactors; do not promote unsupported facts into stable organization capabilities. If an existing source does not provide a trustworthy binding for a specific field, report the limitation rather than inventing one. Expected Deliverables We expect: extension of the existing Python capability/profile layer; implementation of the agreed additional organization capabilities; integration with the existing Ngo model and existing interfaces; preservation of VERIFIED / DOCUMENT_EVIDENCED / DECLARED / UNKNOWN semantics; provenance for positive facts; deterministic automated tests; anti-leakage tests; representative real-document tests; fail-closed handling of ambiguous or missing information; concise technical documentation; a short coverage report describing: supported fields; source used for each field; evidence status; unsupported cases; known limitations. Definition of Done The project will be considered complete when: the agreed additional organization profile fields are supported; existing functionality continues to work; existing organization type capability behavior is preserved; every positive fact has traceable evidence/provenance; VERIFIED, DOCUMENT_EVIDENCED, DECLARED and UNKNOWN information remain clearly distinguishable; unrelated free-text content cannot silently create stable capabilities; ambiguous or unsupported information remains UNKNOWN; automated tests pass; representative real inputs produce the agreed expected outputs; the implementation is delivered as a bounded extension to the existing codebase and is ready for independent review. Out of Scope This project does not include: building the organization intake form; frontend development; n8n orchestration; opportunity/grant document extraction; matching logic redesign; database redesign; building a new organization profile system from scratch; rebuilding the existing platform. Ideal Candidate We are particularly interested in developers with experience in: Python; document intelligence; structured information extraction; NLP; LLM structured outputs; schema-driven data modeling; evidence/provenance systems; PDF processing; automated testing / pytest; working safely inside an existing production-oriented codebase. When Applying Please briefly explain: how you would extend an existing evidence-backed organization profile system without redesigning it; how you would preserve VERIFIED vs DOCUMENT_EVIDENCED vs DECLARED vs UNKNOWN semantics; where you would use deterministic extraction versus LLM-assisted extraction; how you would prevent unrelated free-text data from becoming stable facts; examples of similar Python/document intelligence work; your estimated delivery time; your estimated number of hours for a bounded v0.1 extension. We are looking for someone who can carefully extend an existing, test-driven evidence-based system, not someone proposing a complete rewrite.

Skills required

This job is listed on Freelancer.com. AiZity aggregates listings for discovery only and is not the employer. To bid or apply, use the button in the sidebar.

Similar jobs

Other open projects with overlapping skills and budget type.

More Python jobs →