DDC CWICR Construction Cost Database

Multilingual construction work-item and resource cost database for AI, with Qdrant semantic search and n8n estimating pipelines.

Stars
★ 241
Last updated
1mo ago
License
NOASSERTION
Primary language
HTML

At a glance

How it runs
Self-hosted serviceChat bot
Works with
Portable with changesOpenAI APIClaude API (Partial support)
Cost
Free software; you pay for model usage
Setup effort
Medium · a few setup steps
You'll need
n8n v1.0+ (v2.0+ needs NODES_EXCLUDE=[])Qdrant (cloud or self-hosted)OpenAI API key (text-embedding-3-large)Vision LLM API key (GPT-4 Vision or Gemini 2.0 Flash)DDC Converter RvtExporter.exe for Revit extractionShell / CLINetwork accessLocal filesystem
Typical use
A cost estimator preparing a tender types "100 m² gypsum partition with studs and skim coat" into the Telegram text bot and gets matched rates, a resource breakdown and a total price in one reply.
Not a fit if
  • Commercial users unwilling to buy the separate DDC commercial licence
  • Individuals who do not want to self-host Qdrant and n8n

What does this agent do, and when should you use it?

DDC CWICR (Construction Work Items, Components & Resources) is an open cost database holding 55,719 work items and 27,672 resources across 30 regions and 23 languages, with every record laid out in an 85-field schema covering classification, rate identity, resources, labor, machinery, price variants, aggregates and mass/services. Data ships as XLSX, Parquet and CSV, plus pre-built Qdrant snapshots that already contain OpenAI text-embedding-3-large (3072-dimension) vectors for 11 language collections, so semantic retrieval works without any embedding job of your own. Four n8n workflow JSONs turn the data into working estimators: a Telegram text bot, a photo web form, a universal bot accepting text/photo/PDF, and a CAD (BIM) pipeline that reads Revit/IFC/DWG exports and produces 4D/5D estimates. Costing follows resource-based norms — labor hours, material quantities and machine time multiplied by regional prices — which keeps the technology norms portable while only prices change per market. Licensing is split: data under CC BY-NC 4.0 plus a separate DDC commercial licence, code under Apache-2.0, and you supply your own OpenAI/Gemini keys plus self-hosted Qdrant and n8n.

The repository is organized by source country — folders such as CIS-Russia-GESN-FER-TER/, Asia-China-Dinge/, Europe-Turkey-Birim-Fiyat/ and SouthAmerica-Brazil-SINAPI/ — each holding the national base as parquet and xlsx (native language plus English), a resource catalog as csv/xlsx, a PROVENANCE.md, and a markets/ directory with the base repriced via World Bank PPP and translated for 48 target markets. Qdrant collections are named ddc_{lang}_{city}; a snapshot like DE___DDC_CWICR/DE_BERLIN_workitems_costs_resources_EMBEDDINGS_3072_DDC_CWICR.snapshot loads collection ddc_de_berlin with 55,719 points. Derived tracks such as ddc_au_sydney (sourced from UK_GBP) keep rate_code, resource_code and all norms bytewise identical to the source track, replacing only prices and translatable text. The bundled n8n workflows read a Telegram message, image or PDF, have an LLM extract work items, call OpenAI text-embedding-3-large to embed them, query Qdrant for matching rates, rerank with an LLM, and write HTML/Excel/PDF estimates. The CAD pipeline runs n8n_4_CAD_(BIM)_Cost_Estimation_Pipeline_4D_5D_with_DDC_CWICR.json through RvtExporter.exe to export a Revit model to XLSX, then processes 10 stages (project detection, phase generation, element assignment, work decomposition, vector search, unit mapping, cost calculation, 7.5 CTO validation, aggregation, report generation). All workflows take credentials from a central 🔑 TOKEN node: bot_token, OPENAI_API_KEY, GEMINI_API_KEY, QDRANT_URL, QDRANT_API_KEY.

  1. A cost estimator preparing a tender types "100 m² gypsum partition with studs and skim coat" into the Telegram text bot and gets matched rates, a resource breakdown and a total price in one reply.
  2. A site manager on a renovation job photographs a room; the photo estimator web form has GPT-4 Vision identify room type and elements, estimate dimensions from reference objects, and return an HTML cost report.
  3. A BIM engineer exports an IFC/Revit model and runs stage 4 decomposition so element types become trades (masonry, mortar, plaster) priced through Qdrant, producing a phased 4D/5D HTML and XLS report.
  4. A contractor operating in several countries benchmarks the same rate across Berlin, Toronto, Paris and São Paulo by switching between ddc_de_berlin and ddc_en_toronto collections.
  5. An AI developer building a cost-estimation RAG assistant loads a 3072-dimension Qdrant snapshot instead of generating its own embeddings.
  6. A data analyst takes the ~55 MB Parquet base into an ETL/BI pipeline to study labor hours and material consumption without running any AI service.

How do you install or deploy this agent?

The database itself needs no install — clone the repo to get the tabular data:

git clone https://github.com/datadrivenconstruction/OpenConstructionEstimate-DDC-CWICR.git
cd OpenConstructionEstimate-DDC-CWICR

If you only want the cost data, read the Parquet/XLSX/CSV files directly. For the full AI pipeline you need n8n v1.0+, a Qdrant instance, and OpenAI plus a vision model API key. On n8n 2.0 and later the Execute Command node is disabled by default and must be re-enabled, otherwise the CAD/BIM pipeline will not run:

set NODES_EXCLUDE=[] && npx n8n

Alternatively create C:\Users\YOUR_USER\.n8n\.env containing NODES_EXCLUDE=[]. For Docker, set NODES_EXCLUDE=[] in the container environment. The repository ships the n8n workflow JSON and Qdrant snapshots but no Qdrant installation script — deploy Qdrant separately; a local instance normally listens on http://localhost:6333.

How do you use this agent?

First, upload a snapshot into Qdrant so the collection name matches the language collection used by the workflow:

curl -X POST "http://localhost:6333/collections/ddc_en_toronto/snapshots/upload" \
  -H "Content-Type: multipart/form-data" \
  -F "snapshot=@EN___DDC_CWICR/EN_TORONTO_workitems_costs_resources_EMBEDDINGS_3072_DDC_CWICR.snapshot"

Second, import a workflow in n8n via Workflows → Import from File, for example n8n_1_Telegram_Bot_Cost_Estimates_and_Rate_Finder_TEXT_DDC_CWICR.json.

Third, fill in the 🔑 TOKEN node:

{
  "bot_token": "YOUR_TELEGRAM_BOT_TOKEN",
  "OPENAI_API_KEY": "YOUR_OPENAI_KEY",
  "GEMINI_API_KEY": "YOUR_GEMINI_KEY",
  "QDRANT_URL": "http://localhost:6333",
  "QDRANT_API_KEY": ""
}

Fourth, activate and test: send /start to the Telegram bot, or open the form URL n8n provides. The CAD/BIM pipeline additionally needs path_to_converter and project_file set, with RvtExporter.exe installed. OpenAI GPT-4o is the default model; switch by enabling one of Anthropic Chat Model2, Google Gemini Chat Model, xAI Grok Chat Model1 or DeepSeek Chat Model and disabling the rest. Reports are written to the project folder as project_YYYY-MM-DD.html and project_YYYY-MM-DD.xls.

What are this agent's strengths and limitations?

Pros
  • Resource-based costing separates labor hours, material quantities and machine time from prices, so the same norm set can be repriced per region instead of rewritten.
  • Ships both tabular data (XLSX/Parquet/CSV) for 55,719 work items and Qdrant snapshots with 3072-dimension vectors already computed.
  • Covers 23 languages, 11 Qdrant language collections and 30 country price tracks, so cross-region benchmarking needs no translation work of your own.
  • Four ready-to-import n8n workflows cover text, photo, PDF and Revit/IFC/DWG inputs, and the LLM layer is swappable across OpenAI, Claude, Gemini, Grok and DeepSeek.
Limitations
  • Data is CC BY-NC 4.0 plus a separate DDC commercial licence, so commercial use requires a purchase; GitHub reports the licence as NOASSERTION.
  • The full pipeline depends on external services — OpenAI embeddings, a vision model and a self-hosted Qdrant — so it cannot run fully offline and incurs per-call model costs.
  • On n8n 2.0+ you must manually set NODES_EXCLUDE=[] or the Execute Command node is hidden and the CAD/BIM workflow breaks.
  • The README's derived-track snapshot table is truncated, and heavy resource-level parquet files plus market-edition vector snapshots are stated to arrive only in the next release.

How does this agent compare with similar options?

Key facts side by side with the most closely related agents.

Agent Source review Form / cost Stars Updated Language Full support on
DDC CWICR Construction Cost Database This agent 41 · Major gaps Self-hosted serviceFree + model costs ★ 241 1mo ago HTML OpenAI API
txtai: All-in-One AI Framework 58 · Major gaps Library / SDKFree + model costs ★ 13k 4d ago Python OpenAI API · Claude API
DecisionBox 44 · Major gaps Self-hosted serviceFree + model costs ★ 120 today Go OpenAI API · Claude API
Memvid Portable Agent Memory 46 · Major gaps Library / SDKFree + model costs ★ 17k 2mo ago Rust OpenAI API

How does FollowAgents rate this agent?

FollowAgents source review · FARS-2.1
Major gaps
41/ 100 5-point scale 2.1 / 5
Trust 10/29
Reliability 3/14
Adaptability 8/18
Convention 10/18
Effectiveness 7/13
Verifiability 3/8
Why each dimension lost points
Trust10 / 29 · 1.7/5

Trust: LICENSE cleanly separates DATA (CC BY-NC 4.0) from CODE (Apache-2.0), names per-folder PROVENANCE.md and attribution requirements, so source_attribution earns 3. least_privilege is only 1: the n8n workflows and adapter scripts need OpenAI/Qdrant/Telegram credentials and no minimal-scope guidance is given. user_confirmation is 0: the documented flow sends text/photo and returns an estimate automatically, with no pre-execution confirmation or human review step. data_flow_transparency is 1: README mentions OpenAI embeddings and Qdrant retrieval but gives no end-to-end data-flow or third-party processing statement. sensitive_data_handling is 1: SECURITY.md acknowledges scripts hold user API keys yet explicitly places key-leaking misconfigurations out of scope, pushing the risk to users. dependency_security is 1: workflows pull qdrant/qdrant:latest with no pinning or hash verification; only LFS pointer hashes cover release artefacts. external_effects is 1: Telegram bots send external messages and call paid LLM APIs with no stated quota, rate-limit or side-effect boundary. rollback is 0: no mechanism to undo or recover from executed workflow side effects.

Reliability3 / 14 · 1.1/5

Reliability: self_consistency is 1 - badges claim 30 regions / 12+ countries while the body and LICENSE list nine country folders, so the counts disagree. dependency_availability is 1 - pandas/pyarrow/openai/qdrant-client are unpinned, the Qdrant image is latest, and large datasets depend on Git LFS plus external snapshot downloads. failure_messages is 0 - no error handling, exception messaging or troubleshooting code appears in the supplied files, and the README troubleshooting section is not actually shown.

Adaptability8 / 18 · 2.2/5

Adaptability: audience_and_scenarios is 2 - README targets cost engineers, contractors and AI developers with text/photo/CAD scenarios. capability_boundaries is 1 - no statement of what falls outside the covered regions or of estimation accuracy limits. trigger_precision is 1 - n8n workflows trigger on free-form natural language with no explicit conditions, input validation or refusal policy. environment_fit is 1 - requires n8n 2.0+, Docker, Qdrant, an OpenAI key and Git LFS, with no lighter alternative path.

Convention10 / 18 · 2.8/5

Convention: license is 3 - the dual-licence structure, commercial path, irrevocability note, EU database right and disclaimers are thorough. information_architecture is 2 - README has a TOC, field groups and format tables, but the 85 fields are listed by name only without types or semantics. install_notes is 1 - only git clone and a four-step n8n import, with no dependency install, snapshot restore or credential detail. naming_stability is 1 - version is only v0.1.1 and the licence notes the CIS data moved from CC BY 4.0 to CC BY-NC 4.0, so licensing terms have shifted. examples_and_faq is 2 - Claude Code prompt examples and a cost-calculation example exist, but no FAQ. known_limitations is 1 - only the Section 6 data-quality disclaimer, with no functional or coverage limits. versioning_changelog is 1 - a version badge exists but no CHANGELOG. maintenance_responsibility is 2 - LICENSE and SECURITY.md name Artem Boiko, a German small-business entity, contact email and response timelines, though the publisher is unverified by FollowAgents.

Effectiveness7 / 13 · 2.7/5

Effectiveness: output_usability is 2 - Excel/Parquet/CSV/Qdrant formats with an 85-field schema are directly usable for ETL and RAG. marginal_value is 2 - 55K+ work items, 27K+ resources, 11 languages and precomputed embeddings are genuinely differentiated. cost_benefit is 1 - the functional suite self-reports ~15 minutes and $3-5 in GPT-judge cost per run, and commercial use requires a separate paid licence, so the cost structure is not light for ordinary users.

Verifiability3 / 8 · 1.9/5

Verifiability: claim_traceability is 1 - the 55,719 work items, 27,672 resources and 11-language figures have no accompanying validation file or counting script. cross_source_corroboration is 1 - only LICENSE and README corroborate each other on licensing and folders, while dataset size and region counts disagree across files. fact_inference_separation is 1 - README mixes marketing claims such as '100+ years of methodology' and 'perfect for AI' with factual statements without separating verifiable fact from inference.

Risks and how to mitigate them
  • Not found in source: confirmation before actingTurn on (or add) a confirmation step before it acts, and try it in a sandbox or test environment before real data.
  • Not found in source: rollback or recovery pathBack up first, or work on a git branch or snapshot, so its changes can be undone.
  • Commercial use is restricted: DATA is CC BY-NC 4.0, so any commercial product or paid service needs a separate commercial licence; misuse may breach the terms.
  • Workflows require OpenAI, Qdrant and Telegram credentials, and SECURITY.md explicitly excludes key-leaking misconfigurations from its security scope, so users must isolate secrets themselves.
  • The automation lacks pre-execution confirmation and rollback; bots send external messages and call paid APIs directly, so add human review and quota limits before production use.
  • Dependencies are unpinned (qdrant/qdrant:latest, unlocked pip packages) and large datasets rely on Git LFS plus external snapshot downloads, creating build reproducibility risk.
  • Badges claim 30 regions / 12+ countries while the licence and body list only nine country folders, so coverage claims are inconsistent and should be verified before use.
  • Publisher identity is unverified by FollowAgents; maintenance responsibility and update paths rest only on in-repo documentation and cannot be used to infer reliability or safety.
Evidence confidence: Low Reviewed Sep 29, 2026 Reviewed revision e9e23fbab910
See the full review method →

FAQ

Can I use this commercially?
The data is CC BY-NC 4.0 (free for non-commercial use with attribution) plus a separate DDC commercial licence, so commercial use requires buying that licence; the repository code is Apache-2.0. GitHub reports the licence as NOASSERTION, so check the PROVENANCE.md inside each national folder and the License section before relying on it.
Do I have to use OpenAI?
Embeddings are somewhat bound to OpenAI: the snapshots contain precomputed vectors, but a query still has to be embedded with text-embedding-3-large to search them. The generation and reranking LLM is swappable between Claude Opus 4, Gemini 2.5 Pro, xAI Grok and DeepSeek, and photo analysis can use either Gemini 2.0 Flash or GPT-4 Vision.
What services and keys does the pipeline require?
At minimum n8n (v1.0+), a Qdrant instance, a Telegram Bot Token, an OpenAI API key, and a Qdrant URL plus API key. The universal bot and photo workflow also need a vision API key from Gemini or OpenAI. The CAD/BIM pipeline additionally needs the DDC Converter RvtExporter.exe and a shell where it can be executed.
After import, the CAD/BIM nodes show question marks — why?
That is n8n 2.0+ disabling Execute Command by default. Set the environment variable NODES_EXCLUDE=[] before starting n8n, or write NODES_EXCLUDE=[] into ~/.n8n/.env, or add it to the environment block of your Docker Compose file. After restarting, searching for Execute Command should show the node.
How do I add a new market or language?
Follow the country-track build pipeline: a derived track keeps rate_code, resource_code and the norms (labor hours, machine hours, resource quantities) bytewise identical and replaces only prices and translatable text. Generate the compact catalog under markets/ the same way, and create a new collection named {LANG}_{CITY}_workitems_costs_resources_EMBEDDINGS_3072_DDC_CWICR.
View on GitHub ↗ Install ↓

Related agents