Machine learningCase study notesRAG

From public data to market intelligence: how Valco AI was built

Dubai publishes its property market as open data, updated daily. It is public, and it is not ready to use. The pipeline, the data model, the forecasting models and the conversational layer we built for Valco Properties, and the order we built them in.

Amir Pournasserian · October 8, 2026 · 3 min read

The Dubai Land Department releases its records of property sales, mortgages and gifts, of registered units and of rent contracts, and updates them daily. For a real estate firm that is a gift: the whole market, as it moves. It is also, as it arrives, unusable. Valco AI is the platform we designed and built for Valco Properties in 2024 and 2025 to turn those files into market intelligence and into a conversational assistant for the company’s staff. This note is about the order we built it in, because the order is the lesson.

First, the data

The files are large CSVs with names in Arabic and English, text where a database wants a key, empty values written as the word “null”, and dates written day first. Nothing about that is unusual. It is what public data looks like.

The pipelines read each file as a stream, one line at a time, so the size of a file never matters. Cleaning turns empty strings and the word “null” into real nulls and puts every date into one format. The cleaned records are bulk-loaded into staging tables in SQL Server, with MongoDB alongside for the documents that do not fit a table.

Then the part that most projects skip: set-based SQL turns the staging tables into a model of the market. Lookup tables hold the areas, projects, master projects, buildings, property types and room counts, and the nearest landmark, metro station and mall. Fact tables hold the transactions, the units and the rent contracts, and they share the same lookups, so a building’s sale prices and its rents can finally be set side by side. Until that model exists, every question about the market is a question about file formats.

Second, the models

On the modelled data we ran time-series forecasting, clustering, anomaly detection and predictive scoring, written in Python on PyTorch, TensorFlow and Keras and including recurrent neural networks for the series that warranted them. Each model was measured against a baseline on the firm’s own history before it earned a place. Forecasting where an area’s prices are heading, grouping comparable properties, flagging a transaction that does not fit, and scoring a property as an investment are four different questions; they share the data model and nothing else.

Third, the conversation

Valco’s staff do not want to write SQL. They talk to the platform in plain language: a client’s budget, preferences and timeline, and a question about where to look. The conversational layer is a recommendation system over the market model, not over a pile of documents. It finds the best options for a client and answers with market insight in context, drawn from the forecasts and scores rather than from text that looks similar to the question. That distinction, a conversational interface over a modelled database, is one we now make on every project where the underlying truth is numbers.

Fourth, the outputs people hand over

An analytics dashboard in Blazor presents the market intelligence and the investment scores. The platform also produces the market analysis, segmentation and forecasting reports that staff give to their clients. The dashboard is for the firm; the reports are for the people the firm advises.

What we took from it

Model the data first. The pipelines and the market model were the first half of the project. Every model, every answer and every report runs on them. A project that starts with the model, or with the chatbot, ends up rebuilding this part later under pressure.

Choose the model by measurement. Recurrent networks earned their place on some series and not others. Fashion is not a selection criterion.

Put the conversation over the database, not over documents. When the truth is numeric, retrieval over text produces confident, wrong answers. Retrieval over a model of the data produces answers that can be checked.

Design for the person who hands the output to a client. The reports mattered as much as the dashboard.

The stack was C# and .NET with SQL Server, MongoDB and Blazor on the platform side, and Python with PyTorch, TensorFlow and Keras for the models. Alongside the platform, Valco AI also licensed our DevGuardian AI code review platform for its internal development team. The service this work belongs to is applied machine learning and forecasting.

Working on this?

These are the patterns we use on client work. Tell us what you are building and we will say how they apply.

Book an AI consultation