Marlen Solutions

Data Engineering and Solutions Architecting

I build data pipelines that run in production, map the source systems they pull from, and write the validation that proves the numbers landed right.

Services

Pipeline builds

Ingestion through delivery, built to run unattended. Scheduled, monitored, and documented well enough that someone else can maintain it.

Source-to-target mapping and data modeling

Field-level mapping documents and dimensional models. The spec is a deliverable, not a byproduct of the build.

Validation and reconciliation

Row counts, value-level comparisons, exception reporting. A job that finishes is not the same as a job that loaded the right data.

Requirements translation

Sitting with the people who own the process, then writing something an engineer can build from. Most of what goes wrong on a data project goes wrong here, not in the code.

Work

Port of Portland

Bid tab analysis pipeline

The Port had years of bid tab data and no way to use it. Historical pricing was fragmented across separate files. There was no way to compare line items across projects or see how costs moved over time.

The workbooks were built for bid openings, not for analysis. Bidder and pricing information varied in structure from source to source, naming conventions were inconsistent, and the same pay item appeared under many different descriptions. Most of the work was reshaping and normalizing that into something a reporting layer could sit on.

I built a layered pipeline in Databricks, raw through reporting, with a crosswalk in the middle that maps normalized descriptions to Port cost codes. The mappings live in reference files rather than in the code, so Port staff change how an item maps by editing a file, not by asking for a developer. Power BI reads the output on a schedule.

It runs as one scheduled job in about five minutes and reports what it found each run, including anything that came through unmapped. The loader halts on a bad reference file edit rather than letting it through, so failures are visible instead of silent.

Adding a new bid tab is now: drop the file in the source folder, run the job, check the report. On lookups, the Port's own estimate is over an hour saved per bid.

Capabilities

Production use means I have built and supported it on a real system. Working knowledge means I have built with it and would scope accordingly.

Platforms

Used in production

  • Databricks
  • SQL Server and T-SQL, including schema definition and complex query authoring in SSMS
  • Oracle
  • Power BI

Working knowledge

  • Synapse
  • Snowflake
  • Postgres
Hogan Marhoefer
SUBJECT
Hogan Marhoefer
DESIGNATION
Principal, Marlen Solutions LLC
BASE OF OPS
Portland, Oregon
REGISTRY
Oregon No. 258911594
DOMAINS
Utilities and energy, public sector, insurance and payroll systems
CORE STACK
SQL, Python, dimensional modeling
METHOD
Requirements to mapping to validation
CREDENTIALS
M.S. Applied Data Science for Business
COVERAGE
GL $1M/$2M, E&O $1M, occurrence form
STATUS
Accepting engagements

Engagement

Marlen Solutions LLC is a single-member Oregon LLC, Registry No. 258911594. General liability at $1M per occurrence and $2M aggregate on an occurrence form. Errors and omissions at $1M per claim. Written to public agency requirements.

Invoicing adapts to your process. Purchase orders, contract line items, per-effort breakdowns, or whatever format your AP system expects.

Contact