Marlen Solutions
Data Engineering and Solutions Architecting
I build data pipelines that run in production, map the source systems they pull from, and write the validation that proves the numbers landed right.
Services
Pipeline builds
Ingestion through delivery, built to run unattended. Scheduled, monitored, and documented well enough that someone else can maintain it.
Source-to-target mapping and data modeling
Field-level mapping documents and dimensional models. The spec is a deliverable, not a byproduct of the build.
Validation and reconciliation
Row counts, value-level comparisons, exception reporting. A job that finishes is not the same as a job that loaded the right data.
Requirements translation
Sitting with the people who own the process, then writing something an engineer can build from. Most of what goes wrong on a data project goes wrong here, not in the code.
Work
Port of Portland
Bid tab analysis pipeline
The Port had years of bid tab data and no way to use it. Historical pricing was fragmented across separate files. There was no way to compare line items across projects or see how costs moved over time.
The workbooks were built for bid openings, not for analysis. Bidder and pricing information varied in structure from source to source, naming conventions were inconsistent, and the same pay item appeared under many different descriptions. Most of the work was reshaping and normalizing that into something a reporting layer could sit on.
I built a layered pipeline in Databricks, raw through reporting, with a crosswalk in the middle that maps normalized descriptions to Port cost codes. The mappings live in reference files rather than in the code, so Port staff change how an item maps by editing a file, not by asking for a developer. Power BI reads the output on a schedule.
It runs as one scheduled job in about five minutes and reports what it found each run, including anything that came through unmapped. The loader halts on a bad reference file edit rather than letting it through, so failures are visible instead of silent.
Adding a new bid tab is now: drop the file in the source folder, run the job, check the report. On lookups, the Port's own estimate is over an hour saved per bid.
Capabilities
Production use means I have built and supported it on a real system. Working knowledge means I have built with it and would scope accordingly.
Used in production
- Databricks
- SQL Server and T-SQL, including schema definition and complex query authoring in SSMS
- Oracle
- Power BI
Working knowledge
- Synapse
- Snowflake
- Postgres
- SUBJECT
- Hogan Marhoefer
- DESIGNATION
- Principal, Marlen Solutions LLC
- BASE OF OPS
- Portland, Oregon
- REGISTRY
- Oregon No. 258911594
- DOMAINS
- Utilities and energy, public sector, insurance and payroll systems
- CORE STACK
- SQL, Python, dimensional modeling
- METHOD
- Requirements to mapping to validation
- CREDENTIALS
- M.S. Applied Data Science for Business
- COVERAGE
- GL $1M/$2M, E&O $1M, occurrence form
- STATUS
- Accepting engagements
Engagement
Marlen Solutions LLC is a single-member Oregon LLC, Registry No. 258911594. General liability at $1M per occurrence and $2M aggregate on an occurrence form. Errors and omissions at $1M per claim. Written to public agency requirements.
Invoicing adapts to your process. Purchase orders, contract line items, per-effort breakdowns, or whatever format your AP system expects.
