Household Auto Ownership#

The MTC auto-ownership model predicts how many vehicles each household owns, using five alternatives: zero, one, two, three, or four-or-more vehicles. Its utilities combine household composition and income with characteristics of the home zone.

This walkthrough follows the complete model workflow. We load and sample a typed store, derive the explanatory variables required by the utilities, simulate household choices, inspect a utility trace, and finally calibrate selected alternative-specific constants to match target market shares.

Set up the example#

Informational logging makes the execution of model steps visible without exposing lower-level debugging output. JAX is used later for a compact summary of simulated choices.

import logging

import jax.numpy as jnp
import pandas as pd

import traveler as tv

tv.log_to_stdout(logging.INFO)
[00:00.86] INFO: Logging to stdout set to level 20

Load and sample the MTC population#

The bundled mini dataset contains typed household, person, tour, land-use, and skim tables. It is large enough to exercise chunked JAX calculations while remaining convenient for an executable documentation example.

import mtc

store = mtc.mini()
store
<Store with keys households, persons, skims, land_use, time_periods, tours>

SampleHouseholds selects 100,000 households and carries along their related people and tours. The step returns a new store with rebuilt relationship indexes, leaving the input store unchanged. A fixed default seed makes the sample reproducible.

from traveler.steps import SampleHouseholds

store = SampleHouseholds(n=100_000).run_on(store)
/Users/jpn/Git/traveler/src/traveler/steps/_sample_households.py:262: UserWarning: n (100000) is greater than the number of households in the store, returning all households.
  return self.run_via_polars(store)

Inspect the sampled inputs#

A few rows are enough to establish the structure of the chooser table. Households carry observed vehicle counts and basic demographic fields; the model will write its simulated result to a separate column.

store.households.to_pandas().head()
hhsize num_workers auto_ownership TAZ HHT income home_zone_id _prng_key_
household_id
2717868 2 1 1 25 1 361000 25 15767225188018600536
763899 1 0 1 6 4 59220 6 14249247257885037026
2222791 2 2 2 9 2 197000 9 11490395827286565799
112477 1 1 0 17 6 2200 17 17101579320042102679
370491 3 1 0 21 1 16500 21 5222445121369432989

Person records provide the ages and employment attributes used to derive household counts such as adults, drivers, workers, and children in different age bands.

store.persons.to_pandas().head()
household_id age PNUM sex pemploy pstudent ptype _prng_key_ _household_idx_
person_id
25671 25671 47 1 1 3 3 4 6571566880290666777 3817
25675 25675 27 1 2 3 2 3 4991301075266436201 4066
25678 25678 30 1 2 3 3 4 16083176361397607792 3498
25683 25683 23 1 1 3 3 4 16587057218973660541 4539
25684 25684 52 1 1 3 3 4 16193036017997595926 4710

The store summary shows the complete set of tables that can be resolved while a model expression is evaluated.

store.info()
<mtc.tables.store.Store>
  households = <JaxTable with shape (5000,)>
  persons = <JaxTable with shape (8212,)>
  skims = <Skims with 21 keys>
  land_use = <JaxTable with shape (25,)>
  time_periods = {'EA': 0, 'AM': 1, 'MD': 2, 'PM': 3, 'EV': 4}
  tours = <JaxTable with shape (20000,)>

Derive the explanatory variables#

The raw tables do not yet contain every variable referenced by auto-ownership utility. Four compute steps prepare them in dependency order:

  1. annotate_landuse calculates density measures for each zone.

  2. annotate_persons derives demographic flags and each person’s home zone.

  3. annotate_households aggregates person attributes and derives household segments.

  4. household_value_of_time assigns deterministic random and median value-of-time fields.

from mtc.models.annotate_households import (
    annotate_households,
    household_value_of_time,
)
from mtc.models.annotate_landuse import annotate_landuse
from mtc.models.annotate_persons import annotate_persons

Before those steps run, info() exposes which declared fields are already available and which are still expected from model components.

store.households.info()
store.persons.info()
<JaxTable id_col=household_id>
- household_id         (5000,) int32
- hhsize               (5000,) int8
- num_workers          (5000,) int8
- auto_ownership       (5000,) int8
- TAZ                  (5000,) int32
- HHT                  (5000,) int8
- income               (5000,) int32
- home_zone_id         (5000,) categorical: (25 categories)
- _prng_key_           (5000,) key<fry>
<JaxTable id_col=person_id>
- person_id            (8212,) int32
- household_id         (8212,) int32
- age                  (8212,) int32
- PNUM                 (8212,) int32
- sex                  (8212,) int32
- pemploy              (8212,) int32
- pstudent             (8212,) int32
- ptype                (8212,) int32
- _prng_key_           (8212,) key<fry>
- _household_idx_      (8212,) int32

run_steps passes each returned store to the next step. Keeping that returned value makes the state transition explicit and preserves the original input semantics.

store = store.run_steps(
    annotate_landuse,
    annotate_persons,
    annotate_households,
    household_value_of_time,
)
[00:02.01] INFO: compute_values step annotate_landuse completed in 0:00:00.05
[00:02.29] INFO: compute_values step annotate_persons completed in 0:00:00.28
[00:02.43] INFO: compute_values step annotate_households completed in 0:00:00.14
[00:02.49] INFO: compute_values step household_value_of_time completed in 0:00:00.06

The store records completed step names as workflow metadata. Validation uses this state when deciding which step-produced fields are required.

store.completed_steps
frozenset({'annotate_households',
           'annotate_landuse',
           'annotate_persons',
           'household_value_of_time'})

The household summary now includes the counts, density-related lookups, income segment, and value-of-time data consumed by the auto-ownership utility function.

store.households.info()
<JaxTable id_col=household_id>
- household_id         (5000,) int32
- hhsize               (5000,) int8
- num_workers          (5000,) int8
- auto_ownership       (5000,) int8
- TAZ                  (5000,) int32
- HHT                  (5000,) int8
- income               (5000,) int32
- home_zone_id         (5000,) categorical: (25 categories)
- _prng_key_           (5000,) key<fry>
- income_in_thousands  (5000,) float32
- income_segment       (5000,) int8
- family               (5000,) bool
- non_family           (5000,) bool
- num_drivers          (5000,) int8
- num_adults           (5000,) int8
- num_children         (5000,) int8
- num_young_children   (5000,) int8
- num_children_5_to_15 (5000,) int8
- num_children_16_to_17 (5000,) int8
- num_college_age      (5000,) int8
- num_young_adults     (5000,) int8
- home_is_urban        (5000,) bool
- home_is_rural        (5000,) bool
- median_value_of_time (5000,) float32
- random_vot           (5000,) float32

A final store-level validation checks the derived arrays against their declared shapes and dtypes. Safe coercion is allowed, while unrelated fields that are not needed in this walkthrough may remain absent.

store.validate(verbose=1, coerce="safe", missing_fields="ignore")
Found 2 category groups for validation.
  TIMEPERIOD: ['EA', 'AM', 'MD', 'PM', 'EV'] (ordered=False)
  TAZ: [ 1  2  3  4  5  6  7  8  9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24
 25] (ordered=False)
table households is <mtc.tables.households.Households object at 0x11f18e7b0>
Validating table 'households'...
Field 'household_id' validated as <class 'numpy.int32'>.
Field 'TAZ' validated as <class 'numpy.int32'>.
Field 'HHT' validated as <class 'numpy.int8'>.
Field 'hhsize' validated as <class 'numpy.int8'>.
Field 'income' validated as <class 'numpy.int32'>.
Field 'auto_ownership' validated as <class 'numpy.int8'>.
Field 'random_vot' validated as <class 'numpy.float32'>.
Field 'income_segment' validated as <class 'numpy.int8'>.
Field 'family' validated as <class 'bool'>.
Field 'non_family' validated as <class 'bool'>.
Field 'num_drivers' validated as <class 'numpy.int8'>.
Field 'num_adults' validated as <class 'numpy.int8'>.
Field 'num_children' validated as <class 'numpy.int8'>.
Field 'num_young_children' validated as <class 'numpy.int8'>.
Field 'num_children_5_to_15' validated as <class 'numpy.int8'>.
Field 'num_children_16_to_17' validated as <class 'numpy.int8'>.
Field 'num_college_age' validated as <class 'numpy.int8'>.
Field 'num_young_adults' validated as <class 'numpy.int8'>.
Field 'num_workers' validated as <class 'numpy.int8'>.
table tours is <mtc.tables.tours.Tours object at 0x11e621940>
Validating table 'tours'...
Field 'tour_id' validated as <class 'numpy.int32'>.
Field 'person_id' validated as <class 'numpy.int32'>.
Field 'household_id' validated as <class 'numpy.int32'>.
Field '_household_idx_' validated as <class 'numpy.int32'>.
Field '_person_idx_' validated as <class 'numpy.int32'>.
table persons is <mtc.tables.persons.Persons object at 0x1342bd0d0>
Validating table 'persons'...
Field 'person_id' validated as <class 'numpy.int32'>.
Field 'household_id' validated as <class 'numpy.int32'>.
Field '_household_idx_' validated as <class 'numpy.int32'>.
Field 'age' coerced to <class 'jax.numpy.int8'> without loss of information.
Field 'sex' coerced to <class 'jax.numpy.int8'> without loss of information.
Field 'ptype' coerced to <class 'jax.numpy.int8'> without loss of information.
Field 'pemploy' coerced to <class 'jax.numpy.int8'> without loss of information.
Field 'pstudent' coerced to <class 'jax.numpy.int8'> without loss of information.
Field 'adult' validated as <class 'bool'>.
Field 'male' validated as <class 'bool'>.
Field 'female' validated as <class 'bool'>.
Field 'age_16_to_19' validated as <class 'numpy.bool'>.
Field 'age_16_p' validated as <class 'numpy.bool'>.
table land_use is <mtc.tables.landuse.LandUse object at 0x134282150>
Validating table 'land_use'...
Field 'TAZ' validated as <class 'numpy.int32'>.
Field 'TOTACRE' validated as <class 'numpy.float32'>.
Field 'county_id' validated as <class 'numpy.int8'>.
Field 'TOTHH' validated as <class 'numpy.int32'>.
Field 'TOTPOP' validated as <class 'numpy.int32'>.
Field 'TOTEMP' validated as <class 'numpy.int32'>.
Field 'RESACRE' validated as <class 'numpy.float32'>.
Field 'CIACRE' validated as <class 'numpy.float32'>.
Field 'DISTRICT' validated as <class 'numpy.int8'>.
Field 'SD' validated as <class 'numpy.int8'>.
Field 'AGE0519' validated as <class 'numpy.int32'>.
Field 'RETEMPN' validated as <class 'numpy.int32'>.
Field 'FPSEMPN' validated as <class 'numpy.int32'>.
Field 'HEREMPN' validated as <class 'numpy.int32'>.
Field 'density_index' validated as <class 'numpy.float32'>.
<Store with keys households, persons, skims, land_use, time_periods, tours>

Simulate auto-ownership choices#

The auto_ownership step evaluates utility for all five vehicle-count alternatives, converts those utilities to multinomial-logit probabilities, and draws one alternative using each household’s stable random stream. run_on writes the result to auto_ownership_sim in a returned store.

from mtc.models.auto_ownership import auto_ownership

simulated_store = auto_ownership.run_on(store)

Keeping the observed and simulated columns side by side makes their roles clear: auto_ownership came from the input population, while auto_ownership_sim is this model run’s prediction.

simulated_store.households.to_polars().select(
    "household_id", "auto_ownership", "auto_ownership_sim"
).head(10)
shape: (10, 3)
household_idauto_ownershipauto_ownership_sim
i32i8i8
271786811
76389910
222279124
11247700
37049101
22688000
56849900
2795601
11070600
230719314

Trace the utility calculation#

For model debugging, the utility function can return selected intermediate expressions alongside the alternative utilities. The trace below exposes income and driver indicators, then combines them with each vehicle-count utility for the first few households.

utility_trace = auto_ownership.utility(store, trace=True)

pd.DataFrame(
    {
        **{f"utility_{cars}": values for cars, values in utility_trace.main.items()},
        **utility_trace.trace,
    }
).head()
utility_0 utility_1 utility_2 utility_3 utility_4 income_in_thousands three_drivers two_drivers
0 0.0 0.835464 -5.751639 -153.313644 1.481800 361.000031 False True
1 0.0 1.060974 -10.055627 -19.713758 -3.026466 59.220001 False False
2 0.0 0.666410 -4.723795 -125.842300 1.350800 197.000015 False True
3 0.0 0.556440 -5.401515 -67.820976 -4.297120 2.200000 False False
4 0.0 1.232409 -0.439981 -50.280209 -0.273450 16.500000 False True

Compare simulated and analytic shares#

A simulated run contains random choice variation. The analytic share instead averages each household’s modeled probabilities, giving the model’s expected aggregate distribution without drawing choices. On a large sample the simulated shares should be close, but not identical, to that expectation.

analytic_shares = auto_ownership.analytic_share(store)
simulated_shares = jnp.bincount(
    simulated_store.households["auto_ownership_sim"], length=5
) / len(simulated_store.households)

{"analytic": analytic_shares, "simulated": simulated_shares}
{'analytic': Array([0.26307496, 0.5855463 , 0.02963491, 0.00279747, 0.11894631],      dtype=float32),
 'simulated': Array([0.264 , 0.5864, 0.0268, 0.0034, 0.1194], dtype=float32)}

Calibrate to target shares#

Suppose an external control total says the desired distribution is 10%, 30%, 40%, 10%, and 10% across the five alternatives. We will adjust only the constants for one through four-or-more vehicles; the zero-vehicle utility remains the reference.

target_shares = {0: 0.1, 1: 0.3, 2: 0.4, 3: 0.1, 4: 0.1}
calibration_parameters = [
    "coef_cars1_asc",
    "coef_cars2_asc",
    "coef_cars3_asc",
    "coef_cars4_asc",
]

The calibration objective is the squared difference between analytic and target shares. Its gradient is available through JAX for every parameter; displaying only the four selected components shows the direction available to the optimizer.

auto_ownership.analytic_share_loss(store, target_shares=target_shares)
Array(0.25510776, dtype=float32)
gradient = auto_ownership.analytic_share_loss_grad(store, target_shares=target_shares)

{name: gradient[name] for name in calibration_parameters}
{'coef_cars1_asc': Array(0.07327148, dtype=float32, weak_type=True),
 'coef_cars2_asc': Array(-0.02221743, dtype=float32, weak_type=True),
 'coef_cars3_asc': Array(0.00015584, dtype=float32, weak_type=True),
 'coef_cars4_asc': Array(-0.01873321, dtype=float32, weak_type=True)}

The calibration helper passes that objective and gradient to SciPy. It returns an optimizer result containing a new parameter collection and the final analytic shares; it does not mutate the model’s original parameters.

calibration = auto_ownership.calibrate_analytic_share(
    store,
    target_shares=target_shares,
    calib_params=calibration_parameters,
)

The fitted constants show exactly what the optimizer changed. All other utility coefficients remain at their original values.

pd.DataFrame(
    {
        "before": [auto_ownership.param[name] for name in calibration_parameters],
        "after": [calibration["param"][name] for name in calibration_parameters],
    },
    index=calibration_parameters,
)
before after
coef_cars1_asc 1.1865 1.679031
coef_cars2_asc -1.0846 6.342871
coef_cars3_asc -3.2502 -3.333954
coef_cars4_asc -5.3130 -2.961099

Finally, we compare the fitted analytic shares with the requested controls. A successful calibration should bring the modeled vector close to the target while leaving the reusable auto_ownership step itself unchanged.

{
    "success": calibration.success,
    "target": target_shares,
    "fitted": calibration["shares"],
}
{'success': True,
 'target': {0: 0.1, 1: 0.3, 2: 0.4, 3: 0.1, 4: 0.1},
 'fitted': Array([1.2501623e-01, 3.2500395e-01, 4.2497578e-01, 6.0965403e-06,
        1.2499795e-01], dtype=float32)}

This sequence separates four concerns that are easy to blur together: preparing valid model inputs, computing utility, drawing reproducible choices, and fitting aggregate controls. Traveler uses the same typed store and functional state transitions throughout, so each stage can be inspected or validated independently.