# Adding manual summary statistics to summary network

**URL:** <https://discuss.bayesflow.org/t/adding-manual-summary-statistics-to-summary-network/63>\
**Category:** General\
**Created:** [February 17, 2024, 10:18am UTC](https://discuss.bayesflow.org/t/adding-manual-summary-statistics-to-summary-network/63 "2024-02-17T10:18:32Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![hazhir](https://avatars.discourse-cdn.com/v4/letter/h/edb3f5/32.png) [@hazhir](https://discuss.bayesflow.org/u/hazhir)\
**Post date:** [February 17, 2024, 10:18am UTC](https://discuss.bayesflow.org/t/adding-manual-summary-statistics-to-summary-network/63/1 "2024-02-17T10:18:32Z")

</div>

Hi,  
In some experiments with stochastic ODE models we are finding it take a lot of data and training for the summary networks (either TimeSeriesTransformer or SequenceNetwork) to learn the process and measurement noise parameters, even if they have rather clear and intuitive signatures in the outputs which we can visually inspect. That motivates the idea of augmenting automated summary statistics with generic (for time series) manually crafted ones. What is the best way to do that? SplitNetwork? Is there any example out there for how to do that?

Thanks,  
hazhir

---

<div class="post-metadata">

**Author:** ![KLDivergence](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.bayesflow.org/kldivergence/32/15_2.png) [@KLDivergence](https://discuss.bayesflow.org/u/KLDivergence)\
**Post date:** [February 17, 2024, 1:48pm UTC](https://discuss.bayesflow.org/t/adding-manual-summary-statistics-to-summary-network/63/2 "2024-02-17T13:48:10Z")

</div>

Hi Hazhir,

We are currently working on a generic and sophisticated solution to this problem. For now, you can simply achieve what you want with a simple configurator that returns the manually crafted summary statistics into the `direct_conditions` dictionary key. The raw data stays into the `summary_conditions` key. They will be combined automatically. Here is an example with the toy model from GitHub:

```auto
import numpy as np
import bayesflow as bf

def simulator(theta, n_obs=50, scale=1.0):
    return np.random.default_rng().normal(loc=theta, scale=scale, size=(n_obs, theta.shape[0]))

def prior(D=2, mu=0., sigma=1.0):
    return np.random.default_rng().normal(loc=mu, scale=sigma, size=D)

def configurator(input_dict):
    
    # Example hand crafted statistics: sample average of shape (batch_size, D)
    stats = np.mean(input_dict['sim_dict'], axis=1).astype(np.float32)

    # Raw data will still be processed by the summary network
    raw_data = input_dict['sim_data'].astype(np.float32)

    output_dict = {
        'summary_conditions': raw_data,
        'direct_conditions': stats,
        'parameters': input_dict['prior_draws'].astype(np.float32)
    }
    return output_dict

generative_model = bf.simulation.GenerativeModel(prior, simulator)

# Inspect output
configurator(generative_model(batch_size=3))

# Workflow as usual...

```

Don’t forget to pass your custom configurator to the `Trainer`.

---

<div class="post-metadata">

**Author:** ![hazhir](https://avatars.discourse-cdn.com/v4/letter/h/edb3f5/32.png) [@hazhir](https://discuss.bayesflow.org/u/hazhir)\
**Post date:** [February 17, 2024, 5:13pm UTC](https://discuss.bayesflow.org/t/adding-manual-summary-statistics-to-summary-network/63/3 "2024-02-17T17:13:08Z")

</div>

Great, thanks, and looking forward to your new solution for this problem.

---

<div class="post-metadata">

**Author:** ![Jice](https://avatars.discourse-cdn.com/v4/letter/j/c77e96/32.png) [@Jice](https://discuss.bayesflow.org/u/Jice)\
**Post date:** [February 17, 2024, 8:44pm UTC](https://discuss.bayesflow.org/t/adding-manual-summary-statistics-to-summary-network/63/4 "2024-02-17T20:44:51Z")

</div>

Hi Hazhir,  
I also encountered this situation like yours. In my case, simulation budget is limited due to very expensive forward simulation. What I did is to manually transform the time-series data to frequency-domain data, such as extracting natural frequency from acceleration time-series data. The use of summary statiscs of natural frequency to train the model is very efficient and requires less training data.

---

<div class="post-metadata">

**Author:** ![ali](https://avatars.discourse-cdn.com/v4/letter/a/8c91f0/32.png) [@ali](https://discuss.bayesflow.org/u/ali)\
**Post date:** [August 6, 2024, 10:04pm UTC](https://discuss.bayesflow.org/t/adding-manual-summary-statistics-to-summary-network/63/5 "2024-08-06T22:04:06Z")

</div>

> [@KLDivergence](#):
>
> ```auto
> output_dict = {
> 'summary_conditions': raw_data,
> 'direct_conditions': stats,
> 'parameters': input_dict['prior_draws'].astype(np.float32)
> }
> 
> ```

I have a question on this. Does this mean that the hand-crafted summary statistics (direct\_conditions) are passed to the summary network and learned by the neural net too?

---

<div class="post-metadata">

**Author:** ![marvinschmitt](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.bayesflow.org/marvinschmitt/32/98_2.png) [@marvinschmitt](https://discuss.bayesflow.org/u/marvinschmitt)\
**Post date:** [August 6, 2024, 10:16pm UTC](https://discuss.bayesflow.org/t/adding-manual-summary-statistics-to-summary-network/63/6 "2024-08-06T22:16:55Z")

</div>

Hi,

Direct conditions are **not** passed to the summary network. Only the “summary conditions” go into the summary network. Then, the output of the summary network is concatenated with the direct conditions, and that’s the conditioning input to the normalizing flow.

See this figure for a conceptual overview:

 ![IMG_2369](https://canada1.discourse-cdn.com/flex007/uploads/bayesflow/original/1X/7bd95bc156177b20c45ab117aff76c34e2a09d4b.png)
