<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://www.simulace.info/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=S%C3%A1ra</id>
	<title>Simulace.info - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://www.simulace.info/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=S%C3%A1ra"/>
	<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php/Special:Contributions/S%C3%A1ra"/>
	<updated>2026-07-27T22:21:18Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.31.1</generator>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20986</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20986"/>
		<updated>2021-01-20T21:04:03Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value than they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend approximately 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply would be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stocks always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 trillion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owners, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how much time after appearance of crisis passed, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE because of low inflation it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopps the QE actions. Based on annual change in Fed reserves stock prices will follow with normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factors, such as US governemnt actions, situation in China as there are many investors to US stock market etc.).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|right]]   [[File:two.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
During simulation on the right I observed another crisis, stock owners wealth resulted in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. When there was not another crisis than pandemic crisis for the next decade, the growth of both stock prices and Fed reserves was not that significant. Stock prices reached 47 000 billion dollars and Fed reserves slightly below 20 trillion dollars. We can also see that stock prices stagnated for about 3 years.&lt;br /&gt;
&lt;br /&gt;
Code: [[File:QE_simulation_1st.nlogo]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
For the second round I established agent group banks, which represent private banks, I gave them variable called propensity to loans. They would make loan in rate of 0.9 normal distribution with std 0.1, which means that almost all of new money will go straight into real economy. I had to remove the function which calculates stock prices according to QE policy, as studies find that it is most likely the fact that currently banks don't lend the most of money they get, why stock prices are highly correlated to QE policy. Now I won't investigate stock prices development, as in this case it will mostly influenced by factors, which are out of scope of this simulation. I also assume, that Fed still wants to keep interest rates low to create &amp;quot;debtor favourable enviroment&amp;quot;. Now banks will affect inflation according to how much they lend. Rate how certain amount of new money in economy influences inflation rate I calculated when I compared how it affected inflation rate in the last decade. I know that now banks lend from 8 to 12% and so I can calculate how they will afect inflation rate when giving another percentage of new money into flowing into economy given yearly change of Fed reserves.&lt;br /&gt;
&lt;br /&gt;
[[File: three.png|500px|left|thumb]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
In the second round it did not matter how many times I ran the simulation, the results were repeatedly very similiar. If I suppose that it is more important to deal with crisis and its consequences than to keep inflation down, I always ended up in high inflation rate (aroun 8 - 10%), what should indicate for Fed to restrain money from economy, however economy wasn't recovered yet. So I found myself in a situation, when I could not find the right solution. After this simulation I have to admit, that eventhough almost all of the money went into economy it was at expense of high inflation. As I suppose there is 5 years recovery period from recession, if I simulate shorter period of time inflation would not go so up, but there is no empirical evidence that economy is able to recovery in shorter period if all money go into economy. And even if I set recovery time to 2 years after crisis appears, inflation still goes to 5%.&lt;br /&gt;
&lt;br /&gt;
Code:[[File: 2ND-Round.nlogo]]&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
In the first round of simulations I focused on stock prices and did not used private banks as agents, because I found exact numbers at what rate private banks are now lending new money and because random distribution between 8 and 12% would not make significant difference in results. I focused on stock prices and found out that if situation continues in current way, most likely Fed reserves will grow along with stocks valuations. However it is critical to mention that because of this stock market overheating, prices growth does not reflect actual companies growth in terms of revenue, production growth etc. However QE is considered a way how to get out of recession, it is very possible that it actually generates another bubble for another recession.&lt;br /&gt;
&lt;br /&gt;
In the second round I used private banks and gave them right to decide to some extent how much they want to give to people/companies in form of loans, here I do not think it is appropriate to investigate stock prices as there are many external factors, which will have more power than QE. And under these circumstances there are no empirical evidences, which could help me simulate the future.&lt;br /&gt;
&lt;br /&gt;
After investigating types of stock owners I found out that these richest stock owners are not traders and seek different things in stocks than dividents and trading, thus they do not participate on trades and don't normally switch to stocks of another companies, therefore I did not use them to decide whether to establish new companies etc. As that would be very uncommon because they focus on companies they often own majority of and there are no data which support their pendling on the equity market.&lt;br /&gt;
&lt;br /&gt;
To sum up my study, based on empirical research and current situation, it doesn't look like QE is sustainable solution for solving recessions, as it can lead to another ones. Overvaluating stock prices can be dangerous as dividend don't have to reflect their prices and thus lead to stock prices fall at some point. Also increasing in spending is mostly done by richest households and by companies which shares grow in price. This approach lacks some mechanism which will drive reserves back to its initial state, rising of inflation through QE is very gentle and it bounces back very quickly.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20618</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20618"/>
		<updated>2021-01-19T22:58:42Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend approximately 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopps the QE actions. Based on annual change in Fed reserves stock prices will follow with normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|right]]   [[File:two.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
During simulation on the right I observed another crisis, stock owners wealth resulted in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. When there was not another crisis than pandemic crisis for the next decade, the growth of both stock prices and Fed reserves was not that significant. Stock prices reached 47 000 billion dollars and Fed reserves slightly below 20 trillion dollars. We can also see that stock prices stagnated for about 3 years.&lt;br /&gt;
&lt;br /&gt;
Code: [[File:QE_simulation_1st.nlogo]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
For the second round I established agent group banks, which represent private banks, I gave them variable called propensity to loans. They would make loan in rate of 0.9 normal distribution with std 0.1, which means that almost all of new money will go straight into real economy. I had to remove the function which calculates stock prices according to QE policy, as studies find that it is most likely the fact that currently banks don't lend the most of money they get, why stock prices are highly correlated to QE policy. Now I won't investigate stock prices development, as in this case it will mostly influenced by factors, which are out of scope of this simulation. I also assume, that Fed still wants to keep interest rates low to create debtor favourable enviroment. Now banks will affect inflation according to how much they lend. Rate how certain amount of new money in economy influences inflation rate I calculated when I compared how it affected inflation rate in the last decade. I know that now banks lend from 8 to 12% and so I can calculate how they will afect inflation rate when giving another percentage of new money into economy given yearly change of Fed reserves.&lt;br /&gt;
&lt;br /&gt;
[[File: three.png|500px|left|thumb]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
In the second round it did not matter how many times I ran the simulation. If I suppose that it is more important to deal with crisis and its consequences, I always ended up in high inflation rate (aroun 8 - 10%), what should indicate for Fed to restrain money from economy, however economy wasn't recovered yet. So I found myself in a situation, when I could not find the right solution. After this simulation I have to admit, that eventhough almost all of the money went into economy it was at expense of high inflation. As I suppose there is 5 years recovery period from recession, if I simulate shorter period of time inflation would not go so up, but there is no empirical evidence that economy is able to recovery in shorter period if all money go into economy. And even if I set recovery time to 2 years after crisis appears, inflation still goes to 5%.&lt;br /&gt;
&lt;br /&gt;
Code:[[File: 2ND-Round.nlogo]]&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
In the first round of simulations I focused on stock prices and did not used private banks as agents, because I found exact numbers at what rate private banks are now lending new money and because distribution between 8 and 12% would not make significant difference in results. I focused on stock prices and found out that if situation continues in current way, most likely Fed reserves will grow along with stocks valuations. However it is critical to mention that because of this stock market overheating, prices grow does not reflect actual companies growth in terms of revenue, production etc. However QE is considered a way how to get out of recession, it is very possible that it actually generates another bubble for another recession.&lt;br /&gt;
&lt;br /&gt;
In the second round I used private banks and gave them right to decide to some extent how much they want to give to poeple/companies in form of loans, here I do not think it is appropriate to investigate stock prices as there are many external factors, which will have more power than QE. And under these circumstances there are no empirical evidences, which could help me simulate the future.&lt;br /&gt;
&lt;br /&gt;
After investigating types of stock owners I found out that these richest stock owners are not traders and seek different things in stocks than dividents and trading, thus they do not participate on trades and don't normally switch to stock of another companies, therefore I did not use them to decide whether to establish new companies etc. As that would be very uncommon as they focus on companies they often own majority of and there are no date which support their switching.&lt;br /&gt;
&lt;br /&gt;
To sum up my study, based on empirical research and current situation, it doesn't look like QE is sustainable solution for solving recessions, as it can lead to another ones. Overvaluating stock prices can be dangerous as dividend don't have to reflect their prices and thus lead to selling. Also increases in spending is mostly done by richest household and by companies which shares grow in price. This approach lacks some mechanism which will drive reserves back to its initial state, rising of inflation is very gentle and it bounces back to 0 very quickly.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20615</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20615"/>
		<updated>2021-01-19T22:56:00Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend approximately 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopps the QE actions. Based on annual change in Fed reserves stock prices will follow with normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|right]]   [[File:two.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
During simulation on the right I observed another crisis, stock owners wealth resultet in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. When there was not another crisis than pandemic crisis for the next decade, the growth of both stock prices and Fed reserves was not that significant. Stock prices reached 47 000 billion dollars and Fed reserves slightly below 20 trillion dollars. We can also see that stock prices stagnated for about 3 years.&lt;br /&gt;
&lt;br /&gt;
Code: [[File:QE_simulation_1st.nlogo]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
For the second round I established agent group banks, which represent private banks, I gave them variable called propensity to loans. They would make loan in rate of 0.9 normal distribution with std 0.1, which means that almost all of new money will go straight into real economy. I had to remove the function which calculates stock prices according to QE policy, as studies find that it is most likely the fact that currently banks don't lend the most of money they get, why stock prices are highly correlated to QE policy. Now I won't investigate stock prices development, as in this case it will mostly influenced by factors, which are out of scope of this simulation. I also assume, that Fed still wants to keep interest rates low to create debtor favourable enviroment. Now banks will affect inflation according to how much they lend. Rate how certain amount of new money in economy influences inflation rate I calculated when I compared how it affected inflation rate in the last decade. I know that now banks lend from 8 to 12% and so I can calculate how they will afect inflation rate when giving another percentage of new money into economy given yearly change of Fed reserves.&lt;br /&gt;
&lt;br /&gt;
[[File: three.png|500px|left|thumb]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
In the second round it did not matter how many times I ran the simulation. If I suppose that it is more important to deal with crisis and its consequences, I always ended up in high inflation rate (aroun 8 - 10%), what should indicate for Fed to restrain money from economy, however economy wasn't recovered yet. So I found myself in a situation, when I could not find the right solution. After this simulation I have to admit, that eventhough almost all of the money went into economy it was at expense of high inflation. As I suppose there is 5 years recovery period from recession, if I simulate shorter period of time inflation would not go so up, but there is no empirical evidence that economy is able to recovery in shorter period if all money go into economy. And even if I set recovery time to 2 years after crisis appears, inflation still goes to 5%.&lt;br /&gt;
&lt;br /&gt;
Code:[[File: 2ND-Round.nlogo]]&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
In the first round of simulations I focused on stock prices and did not used private banks as agents, because I found exact numbers at what rate private banks are now lending new money and because distribution between 8 and 12% would not make significant difference in results. I focused on stock prices and found out that if situation continues in current way, most likely Fed reserves will grow along with stocks valuations. However it is critical to mention that because of this stock market overheating, prices grow does not reflect actual companies growth in terms of revenue, production etc. However QE is considered a way how to get out of recession, it is very possible that it actually generates another bubble for another recession.&lt;br /&gt;
&lt;br /&gt;
In the second round I used private banks and gave them right to decide to some extent how much they want to give to poeple/companies in form of loans, here I do not think it is appropriate to investigate stock prices as there are many external factors, which will have more power than QE. And under these circumstances there are no empirical evidences, which could help me simulate the future.&lt;br /&gt;
&lt;br /&gt;
After investigating types of stock owners I found out that these richest stock owners are not traders and seek different things in stocks than dividents and trading, thus they do not participate on trades and don't normally switch to stock of another companies, therefore I did not use them to decide whether to establish new companies etc. As that would be very uncommon as they focus on companies they often own majority of and there are no date which support their switching.&lt;br /&gt;
&lt;br /&gt;
To sum up my study, based on empirical research and current situation, it doesn't look like QE is sustainable solution for solving recessions, as it can lead to another ones. Overvaluating stock prices can be dangerous as dividend don't have to reflect their prices and thus lead to selling. Also increases in spending is mostly done by richest household and by companies which shares grow in price. This approach lacks some mechanism which will drive reserves back to its initial state, rising of inflation is very gentle and it bounces back to 0 very quickly.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20612</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20612"/>
		<updated>2021-01-19T22:52:09Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend approximately 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|right]]   [[File:two.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
During simulation on the right I observed another crisis, stock owners wealth resultet in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. When there was not another crisis than pandemic crisis for the next decade, the growth of both stock prices and Fed reserves was not that significant. Stock prices reached 47 000 billion dollars and Fed reserves slightly below 20 trillion dollars. We can also see that stock prices stagnated for about 3 years.&lt;br /&gt;
&lt;br /&gt;
Code: [[File:QE_simulation_1st.nlogo]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
For the second round I established agent group banks, which represent private banks, I gave them variable called propensity to loans. They would make loan in rate of 0.9 normal distribution with std 0.1, which means that almost all of new money will go straight into real economy. I had to remove the function which calculates stock prices according to QE policy, as studies find that it is most likely the fact that currently banks don't lend the most of money they get, why stock prices are highly correlated to QE policy. Now I won't investigate stock prices development, as in this case it will mostly influenced by factors, which are out of scope of this simulation. I also assume, that Fed still wants to keep interest rates low to create debtor favourable enviroment. Now banks will affect inflation according to how much they lend. Rate how certain amount of new money in economy influences inflation rate I calculated when I compared how it affected inflation rate in the last decade. I know that now banks lend from 8 to 12% and so I can calculate how they will afect inflation rate when giving another percentage of new money into economy given yearly change of Fed reserves.&lt;br /&gt;
&lt;br /&gt;
[[File: three.png|500px|left|thumb]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
In the second round it did not matter how many times I ran the simulation. If I suppose that it is more important to deal with crisis and its consequences, I always ended up in high inflation rate (aroun 8 - 10%), what should indicate for Fed to restrain money from economy, however economy wasn't recovered yet. So I found myself in a situation, when I could not find the right solution. After this simulation I have to admit, that eventhough almost all of the money went into economy it was at expense of high inflation. As I suppose there is 5 years recovery period from recession, if I simulate shorter period of time inflation would not go so up, but there is no empirical evidence that economy is able to recovery in shorter period if all money go into economy. And even if I set recovery time to 2 years after crisis appears, inflation still goes to 5%.&lt;br /&gt;
&lt;br /&gt;
Code:[[File: 2ND-Round.nlogo]]&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
In the first round of simulations I focused on stock prices and did not used private banks as agents, because I found exact numbers at what rate private banks are now lending new money and because distribution between 8 and 12% would not make significant difference in results. I focused on stock prices and found out that if situation continues in current way, most likely Fed reserves will grow along with stocks valuations. However it is critical to mention that because of this stock market overheating, prices grow does not reflect actual companies growth in terms of revenue, production etc. However QE is considered a way how to get out of recession, it is very possible that it actually generates another bubble for another recession.&lt;br /&gt;
&lt;br /&gt;
In the second round I used private banks and gave them right to decide to some extent how much they want to give to poeple/companies in form of loans, here I do not think it is appropriate to investigate stock prices as there are many external factors, which will have more power than QE. And under these circumstances there are no empirical evidences, which could help me simulate the future.&lt;br /&gt;
&lt;br /&gt;
After investigating types of stock owners I found out that these richest stock owners are not traders and seek different things in stocks than dividents and trading, thus they do not participate on trades and don't normally switch to stock of another companies, therefore I did not use them to decide whether to establish new companies etc. As that would be very uncommon as they focus on companies they often own majority of and there are no date which support their switching.&lt;br /&gt;
&lt;br /&gt;
To sum up my study, based on empirical research and current situation, it doesn't look like QE is sustainable solution for solving recessions, as it can lead to another ones. Overvaluating stock prices can be dangerous as dividend don't have to reflect their prices and thus lead to selling. Also increases in spending is mostly done by richest household and by companies which shares grow in price. This approach lacks some mechanism which will drive reserves back to its initial state, rising of inflation is very gentle and it bounces back to 0 very quickly.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20608</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20608"/>
		<updated>2021-01-19T22:37:17Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|right]]   [[File:two.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
During simulation on the right I observed another crisis, stock owners wealth resultet in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. When there was not another crisis than pandemic crisis for the next decade, the growth of both stock prices and Fed reserves was not that significant. Stock prices reached 47 000 billion dollars and Fed reserves slightly below 20 trillion dollars. We can also see that stock prices stagnated for about 3 years.&lt;br /&gt;
&lt;br /&gt;
Code: [[File:QE_simulation_1st.nlogo]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
For the second round I established agent group banks, which represent private banks, I gave them variable called propensity to loans. They would make loan in rate of 0.9 normal distribution with std 0.1, which means that almost all of new money will go straight into real economy. I had to remove the function which calculates stock prices according to QE policy, as studies find that it is most likely the fact that currently banks don't lend the most of money they get, why stock prices are highly correlated to QE policy. Now I won't investigate stock prices development, as in this case it will mostly influenced by factors, which are out of scope of this simulation. I also assume, that Fed still wants to keep interest rates low to create debtor favourable enviroment. Now banks will affect inflation according to how much they lend. Rate how certain amount of new money in economy influences inflation rate I calculated when I compared how it affected inflation rate in the last decade. I know that now banks lend from 8 to 12% and so I can calculate how they will afect inflation rate when giving another percentage of new money into economy given yearly change of Fed reserves.&lt;br /&gt;
&lt;br /&gt;
[[File: three.png|500px|left|thumb]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
In the second round it did not matter how many times I ran the simulation. If I suppose that it is more important to deal with crisis and its consequences, I always ended up in high inflation rate (aroun 8 - 10%), what should indicate for Fed to restrain money from economy, however economy wasn't recovered yet. So I found myself in a situation, when I could not find the right solution. After this simulation I have to admit, that eventhough almost all of the money went into economy it was at expense of high inflation. As I suppose there is 5 years recovery period from recession, if I simulate shorter period of time inflation would not go so up, but there is no empirical evidence that economy is able to recovery in shorter period if all money go into economy. And even if I set recovery time to 2 years after crisis appears, inflation still goes to 5%.&lt;br /&gt;
&lt;br /&gt;
Code:[[File: 2ND-Round.nlogo]]&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
In the first round of simulations I focused on stock prices and did not used private banks as agents, because I found exact numbers at what rate private banks are now lending new money and because distribution between 8 and 12% would not make significant difference in results. I focused on stock prices and found out that if situation continues in current way, most likely Fed reserves will grow along with stocks valuations. However it is critical to mention that because of this stock market overheating, prices grow does not reflect actual companies growth in terms of revenue, production etc. However QE is considered a way how to get out of recession, it is very possible that it actually generates another bubble for another recession.&lt;br /&gt;
&lt;br /&gt;
In the second round I used private banks and gave them right to decide to some extent how much they want to give to poeple/companies in form of loans, here I do not think it is appropriate to investigate stock prices as there are many external factors, which will have more power than QE. And under these circumstances there are no empirical evidences, which could help me simulate the future.&lt;br /&gt;
&lt;br /&gt;
After investigating types of stock owners I found out that these richest stock owners are not traders and seek different things in stocks than dividents and trading, thus they do not participate on trades and don't normally switch to stock of another companies, therefore I did not use them to decide whether to establish new companies etc. As that would be very uncommon as they focus on companies they often own majority of and there are no date which support their switching.&lt;br /&gt;
&lt;br /&gt;
To sum up my study, based on empirical research and current situation, it doesn't look like QE is sustainable solution for solving recessions, as it can lead to another ones. Overvaluating stock prices can be dangerous as dividend don't have to reflect their prices and thus lead to selling. Also increases in spending is mostly done by richest household and by companies which shares grow in price. This approach lacks some mechanism which will drive reserves back to its initial state, rising of inflation is very gentle and it bounces back to 0 very quickly.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=File:2ND-Round.nlogo&amp;diff=20607</id>
		<title>File:2ND-Round.nlogo</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=File:2ND-Round.nlogo&amp;diff=20607"/>
		<updated>2021-01-19T22:36:16Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=File:Three.png&amp;diff=20606</id>
		<title>File:Three.png</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=File:Three.png&amp;diff=20606"/>
		<updated>2021-01-19T22:27:37Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20605</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20605"/>
		<updated>2021-01-19T22:27:17Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|right]]   [[File:two.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
During simulation on the right I observed another crisis, stock owners wealth resultet in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. When there was not another crisis than pandemic crisis for the next decade, the growth of both stock prices and Fed reserves was not that significant. Stock prices reached 47 000 billion dollars and Fed reserves slightly below 20 trillion dollars. We can also see that stock prices stagnated for about 3 years.&lt;br /&gt;
&lt;br /&gt;
Code: [[File:QE_simulation_1st.nlogo]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
For the second round I established agent group banks, which represent private banks, I gave them variable called propensity to loans. They would make loan in rate of 0.9 normal distribution with std 0.1, which means that almost all of new money will go straight into real economy. I had to remove the function which calculates stock prices according to QE policy, as studies find that it is most likely the fact that currently banks don't lend the most of money they get, why stock prices are highly correlated to QE policy. Now I won't investigate stock prices development, as in this case it will mostly influenced by factors, which are out of scope of this simulation. I also assume, that Fed still wants to keep interest rates low to create debtor favourable enviroment. Now banks will affect inflation according to how much they lend. Rate how certain amount of new money in economy influences inflation rate I calculated when I compared how it affected inflation rate in the last decade. I know that now banks lend from 8 to 12% and so I can calculate how they will afect inflation rate when giving another percentage of new money into economy given yearly change of Fed reserves.&lt;br /&gt;
&lt;br /&gt;
[[File: three.png|500px|left|thumb]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
In the first round of simulations I focused on stock prices and did not used private banks as agents, because I found exact numbers at what rate private banks are now lending new money and because distribution between 8 and 12% would not make significant difference in results. I focused on stock prices and found out that if situation continues in current way, most likely Fed reserves will grow along with stocks valuations. However it is critical to mention that because of this stock market overheating, prices grow does not reflect actual companies growth in terms of revenue, production etc. However QE is considered a way how to get out of recession, it is very possible that it actually generates another bubble for another recession.&lt;br /&gt;
&lt;br /&gt;
In the second round I used private banks and gave them right to decide to some extent how much they want to give to poeple/companies in form of loans, here I do not think it is appropriate to investigate stock prices as there are many external factors, which will have more power than QE. And under these circumstances there are no empirical evidences, which could help me simulate the future.&lt;br /&gt;
&lt;br /&gt;
After investigating types of stock owners I found out that these richest stock owners are not traders and seek different things in stocks than dividents and trading, thus they do not participate on trades and don't normally switch to stock of another companies, therefore I did not use them to decide whether to establish new companies etc. As that would be very uncommon as they focus on companies they often own majority of and there are no date which support their switching.&lt;br /&gt;
&lt;br /&gt;
To sum up my study, based on empirical research and current situation, it doesn't look like QE is sustainable solution for solving recessions, as it can lead to another ones. Overvaluating stock prices can be dangerous as dividend don't have to reflect their prices and thus lead to selling. Also increases in spending is mostly done by richest household and by companies which shares grow in price. This approach lacks some mechanism which will drive reserves back to its initial state, rising of inflation is very gentle and it bounces back to 0 very quickly.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20604</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20604"/>
		<updated>2021-01-19T22:20:20Z</updated>

		<summary type="html">&lt;p&gt;Sára: /* Conclusion */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|right]]   [[File:two.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
During simulation on the right I observed another crisis, stock owners wealth resultet in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. When there was not another crisis than pandemic crisis for the next decade, the growth of both stock prices and Fed reserves was not that significant. Stock prices reached 47 000 billion dollars and Fed reserves slightly below 20 trillion dollars. We can also see that stock prices stagnated for about 3 years.&lt;br /&gt;
&lt;br /&gt;
Code: [[File:QE_simulation_1st.nlogo]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
For the second round I established agent group banks, which represent private banks, I gave them variable called propensity to loans. They would make loan in rate of 0.9 normal distribution with std 0.1, which means that almost all of new money will go straight into real economy. I had to remove the function which calculates stock prices according to QE policy, as studies find that it is most likely the fact that currently banks don't lend the most of money they get, why stock prices are highly correlated to QE policy. Now I won't investigate stock prices development, as in this case it will mostly influenced by factors, which are out of scope of this simulation. I also assume, that Fed still wants to keep interest rates low to create debtor favourable enviroment. Now banks will affect inflation according to how much they lend. Rate how certain amount of new money in economy influences inflation rate I calculated when I compared how it affected inflation rate in the last decade. I know that now banks lend from 8 to 12% and so I can calculate how they will afect inflation rate when giving another percentage of new money into economy given yearly change of Fed reserves.&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
In the first round of simulations I focused on stock prices and did not used private banks as agents, because I found exact numbers at what rate private banks are now lending new money and because distribution between 8 and 12% would not make significant difference in results. I focused on stock prices and found out that if situation continues in current way, most likely Fed reserves will grow along with stocks valuations. However it is critical to mention that because of this stock market overheating, prices grow does not reflect actual companies growth in terms of revenue, production etc. However QE is considered a way how to get out of recession, it is very possible that it actually generates another bubble for another recession.&lt;br /&gt;
&lt;br /&gt;
In the second round I used private banks and gave them right to decide to some extent how much they want to give to poeple/companies in form of loans, here I do not think it is appropriate to investigate stock prices as there are many external factors, which will have more power than QE. And under these circumstances there are no empirical evidences, which could help me simulate the future.&lt;br /&gt;
&lt;br /&gt;
After investigating types of stock owners I found out that these richest stock owners are not traders and seek different things in stocks than dividents and trading, thus they do not participate on trades and don't normally switch to stock of another companies, therefore I did not use them to decide whether to establish new companies etc. As that would be very uncommon as they focus on companies they often own majority of and there are no date which support their switching.&lt;br /&gt;
&lt;br /&gt;
To sum up my study, based on empirical research and current situation, it doesn't look like QE is sustainable solution for solving recessions, as it can lead to another ones. Overvaluating stock prices can be dangerous as dividend don't have to reflect their prices and thus lead to selling. Also increases in spending is mostly done by richest household and by companies which shares grow in price. This approach lacks some mechanism which will drive reserves back to its initial state, rising of inflation is very gentle and it bounces back to 0 very quickly.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20599</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20599"/>
		<updated>2021-01-19T21:56:16Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|right]]   [[File:two.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
During simulation on the right I observed another crisis, stock owners wealth resultet in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. When there was not another crisis than pandemic crisis for the next decade, the growth of both stock prices and Fed reserves was not that significant. Stock prices reached 47 000 billion dollars and Fed reserves slightly below 20 trillion dollars. We can also see that stock prices stagnated for about 3 years.&lt;br /&gt;
&lt;br /&gt;
Code: [[File:QE_simulation_1st.nlogo]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
For the second round I established agent group banks, which represent private banks, I gave them variable called propensity to loans. They would make loan in rate of 0.9 normal distribution with std 0.1, which means that almost all of new money will go straight into real economy. I had to remove the function which calculates stock prices according to QE policy, as studies find that it is most likely the fact that currently banks don't lend the most of money they get, why stock prices are highly correlated to QE policy. Now I won't investigate stock prices development, as in this case it will mostly influenced by factors, which are out of scope of this simulation. I also assume, that Fed still wants to keep interest rates low to create debtor favourable enviroment. Now banks will affect inflation according to how much they lend. Rate how certain amount of new money in economy influences inflation rate I calculated when I compared how it affected inflation rate in the last decade. I know that now banks lend from 8 to 12% and so I can calculate how they will afect inflation rate when giving another percentage of new money into economy given yearly change of Fed reserves.&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=File:QE_simulation_1st.nlogo&amp;diff=20583</id>
		<title>File:QE simulation 1st.nlogo</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=File:QE_simulation_1st.nlogo&amp;diff=20583"/>
		<updated>2021-01-19T21:10:15Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20582</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20582"/>
		<updated>2021-01-19T21:09:58Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|right]]   [[File:two.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
During simulation on the right I observed another crisis, stock owners wealth resultet in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. When there was not another crisis than pandemic crisis for the next decade, the growth of both stock prices and Fed reserves was not that significant. Stock prices reached 47 000 billion dollars and Fed reserves slightly below 20 trillion dollars. We can also see that stock prices stagnated for about 3 years.&lt;br /&gt;
&lt;br /&gt;
Code: [[File:QE_simulation_1st.nlogo]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20579</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20579"/>
		<updated>2021-01-19T21:04:45Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|right]]   [[File:two.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
During simulation on the right I observed another crisis, stock owners wealth resultet in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. When there was not another crisis than pandemic crisis for the next decade, the growth of both stock prices and Fed reserves was not that significant. Stock prices reached 47 000 billion dollars and Fed reserves slightly below 20 trillion dollars. We can also see that stock prices stagnated for about 3 years.&lt;br /&gt;
&lt;br /&gt;
Code: [[File:QE_simulation.nlogo]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=File:Two.png&amp;diff=20576</id>
		<title>File:Two.png</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=File:Two.png&amp;diff=20576"/>
		<updated>2021-01-19T20:55:50Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20575</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20575"/>
		<updated>2021-01-19T20:53:48Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|caption|left]]&lt;br /&gt;
&lt;br /&gt;
During this simulation I observed another crisis, stock owners wealth resultet in 61 800 billion dollars, Fed reserved reached 20 trillion dollars. &lt;br /&gt;
&lt;br /&gt;
[[File:two.png]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20573</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20573"/>
		<updated>2021-01-19T20:50:41Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|caption|left]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20570</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20570"/>
		<updated>2021-01-19T20:47:26Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|thumb|left]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=File:One.png&amp;diff=20568</id>
		<title>File:One.png</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=File:One.png&amp;diff=20568"/>
		<updated>2021-01-19T20:45:13Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20567</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20567"/>
		<updated>2021-01-19T20:44:30Z</updated>

		<summary type="html">&lt;p&gt;Sára: /* 1st Round */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
[[File:one.png|500px|left]]&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=WS_2020/2021&amp;diff=20565</id>
		<title>WS 2020/2021</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=WS_2020/2021&amp;diff=20565"/>
		<updated>2021-01-19T20:41:33Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Semestral papers from winter term 2020/2021. Please, put here links to the pages with your paper. First you need to have your [[Assignments WS 2020/2021|assignment approved]]&lt;br /&gt;
&lt;br /&gt;
[http://www.simulace.info/index.php/Impact_of_late_lockdown_and_hospital_capacity_on_19-COVID_spread Impact of late lockdown and hospital capacity on 19-COVID spread] , Thomas BAEUMLIN [[User:Toscool|Toscool]] ([[User talk:Toscool|talk]]) 16:15, 7 January 2020 (CET)&lt;br /&gt;
&lt;br /&gt;
[[Comparing_multiple_airplane_boarding_methods]], Adam Jedlička [[User:Jeda00|jeda00]] 20:45, 17 January 2020 (CET)&lt;br /&gt;
&lt;br /&gt;
[http://www.simulace.info/index.php/Effects_of_VAT_exemption_abolition_to_imported_low_value_consignments Effects of VAT exemption abolition to imported low value consignments], Adela Barchankova [[User:bara15|bara15]] ([[User talk:bara15|talk]]) 21:00, 18 January 2020 (CET)&lt;br /&gt;
 &lt;br /&gt;
[[Retirement Planning]], [[User:Louis|Louis]] ([[User talk:Louis|talk]]) 02:08, 19 January 2021 (CET)&lt;br /&gt;
&lt;br /&gt;
[[Covid 19 - Contacts]], [[User:praa03|praa03]] 17:08, 19 January 2021 (CET)&lt;br /&gt;
&lt;br /&gt;
[[Corrupted Blood Incident]], [[User:Frym06|Frym06]] ([[User talk:Frym06|talk]]) 19:32, 19 January 2021 (CET)&lt;br /&gt;
&lt;br /&gt;
WIP [[Car acquisition]], [[Mico00|Mico00]] 19:59, 19 January 2021 (CET)&lt;br /&gt;
&lt;br /&gt;
[[Widening the spread between rich and poor]], [[User:Sára|Sára]] 21:40, 19 January 2021 (CET)&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20561</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20561"/>
		<updated>2021-01-19T20:38:02Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow-minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries, of different kinds and maturities, here we can find one difference from traditional expansionary policy, which aims on federal funds rates) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher and money supply will be also very high. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of deflation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:100 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 trillion dollars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when for appreciating or depreciating their shares, as they probably own shares of different companies and thus different shares can follow different trends.&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance - Initial state 7.1 trillion dollars. I don't actually care for asset balance on FED's accounts (but I use it so that I can obtain also infor about estimated banks asset balance in 10 years), I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
In the first round I use only two types of agents, central bank and stock owners. For simplicity I give same wealth to each of 1000 stock owners. I then simulate Fed actions, in case of new crisis they will do massive QE, in case we are in post-crisis area (1-5 years after crisis), Fed applies QE in smaller amounts than right after crisis, but still in significant rate. Then Fed needs to re-establish QE, if inflation is at or below 0%, but also in smaller power. Otherwise Fed will keep their current reserves rate or gently decrease the rate. Inflation and interest rates react to this change. In case of QE inflation rises (inflation rises weakly due to poor bank lending, cca 4-12% of received capital), but if QE this year doesn't conduct QE, inflation lowers rather more (this observation I got from comparing QE activities in the last decade to inflation changes). Also when QE is present I always assume that interest rates are at 0, or closely to zero. Lastly stock prizes react. In case of Fed reserves rising, there is very small probability (1:500), that stock prizes don't follow, otherwise if QE stagnates, there is 1:10 probability that stock prices won't follow. (There are many reasons for this, as QE is not the only factor, that influences stock prices, apparently it is major factor when QE is present, probably because most of new created money goes into stock market and acts there as inflation in economy, thus rising prices. Other way how we can explain is, that when QE is present, loans are very cheap for companies also and so they can finance dividends using loans, because of almost zero interest rates, thus become more popular than other ways of investment and thus there is higher demand for them. On the other side if QE is not present, stock prices are influenced by other factor, such as US governemnt actions, situation in China as there are many investors to US stock market).&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20419</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20419"/>
		<updated>2021-01-19T12:42:32Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of delation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches below 0.5% or interest rates get above 2.5% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:1000 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 billion dillars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when deciding whether to buy new stocks, or sell. &lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 3% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance_change - I don't actually care for asset balance on FED's accounts, I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;br /&gt;
&lt;br /&gt;
==1st Round==&lt;br /&gt;
&lt;br /&gt;
==2nd Round==&lt;br /&gt;
&lt;br /&gt;
=Conclusion=&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=b0ab3804-b496-4672-85d2-6fc1a1f1caf8%40sessionmgr101&lt;br /&gt;
&lt;br /&gt;
http://web.b.ebscohost.com/ehost/pdfviewer/pdfviewer?vid=0&amp;amp;sid=ed7f3ec8-7c2b-4a07-8f7c-6cc54d92e27c%40pdc-v-sessmgr01&lt;br /&gt;
&lt;br /&gt;
https://search.proquest.com/docview/1039275454/C82BA7C8A9474B3FPQ/1?accountid=17203&lt;br /&gt;
&lt;br /&gt;
https://www.thebalance.com/what-is-quantitative-easing-definition-and-explanation-3305881&lt;br /&gt;
&lt;br /&gt;
https://positivemoney.org/how-money-works/advanced/how-quantitative-easing-works/&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20418</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20418"/>
		<updated>2021-01-19T09:58:34Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of delation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches above 2.5% or interest rates get above 2.5% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. Each year will represent one round in the simulation, there will be also probability of 1:1000 that something unexpected happens and pulls economy into recession.&lt;br /&gt;
&lt;br /&gt;
Before QE the valuation of US equity market was about 12 billion dollars. Now it is around 50 billion dillars [https://siblisresearch.com/data/us-stock-market-value/]. For purpose of this simulation I assume that 5% of richest people own 40% of this market, what is currently 20 323 billion dollars. This is also my target group for investigation. I create group of 1000 stock owner, each of them will own 20.3 billions dollars. I will use them when deciding whether to buy new stocks, or sell. &lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
Time after crisis - this variable will hold the information about how long after crisis it is, if it is more than 5 years Fed will keep its reserves, slowly declining it, buying new assests only after old ones expire. When an crisis occur the number will be minimized to 0 (I will assume that crisis will appear either after some unexpected news, when applying QE only because of low inflation and interest rates it won't be done in 5 year interval but rather until inflation is up at about 1.5% again). Variable starts at 0, because 2020 is the start of pandemic crisis.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
Balance_change - I don't actually care for asset balance on FED's accounts, I need annual change to be able to predict stock prices change. As studies shown stocks prices changes are dependend on Fed policy, but there also other external factors that influence its development (in USA it was Donald Trump's strategy towards domestic producers for example). On the other hand studies show that when Fed's reserves were raising, stock were always raising and when QE stopped stock prices either stopped or raised or slumped a little bit. So I will assume that there is only 1:500 probability that stock don't raise, when Fed applies QE, but 1:10 stock prices don't follow, when Fed stopped the QE actions. Based on annual change in Fed reserves stock prices will follow with log-normal distribution with mean same as QE change and std 10.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20385</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20385"/>
		<updated>2021-01-17T20:25:49Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds and mortgage backed securities, in case of US bonds are US Treasuries) from banks in exchange for new money. The thought behind quantitative easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable for banks to lend money to people, as with low interest rates and present inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after quantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will observe inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
I start in year 2020. We can observe that quantitative easing takes place when interest rates reaches almost 0, but inflation is still very low, and there is threat of economy going into deflation spiral. After global crisis which ended in 2009, Fed used open market operations to lower interest rates, however it did not promote consumption and production enough to get US out of crisis (possibly because americans were already indebted enough and did not want to take new loans), this year there was also deflation of 0.36%. Fed then established quantitative easing, in order to promote spending and get USA out of delation and out of crisis. As of 2020, inflation started to lower again as a result of covid crisis and thus Fed started with massive quantitative easing again. As from last decade we observe that Fed used quantitative easing further after getting country out of crisis to promote economy reviving. This process went for about 5 years after crisis, so I will assume that after covid crisis rounds of quantitative easing will continue. If the inflation reaches above 2.5% or interest rates get above 2.5% Fed will re-eastablish quantitative easing (purchasing US Treasuries / mortgage backed securities) from private banks. As studies suggest, QE is strongly correlated to rise in stock prices (I expect this trend will not be applicable after I tweak rate of lending providen by private banks), however QE is of course not the only variable influencing stock prices. As when Fed balance was increasing, prices of stock always increased as well, but when Fed stopped purchasing and started to slowly decrease its balance, stock prices were still increasing (at that time it was due to Donald Trump elections and his policy of tax reliefs and domestic production promotion). These factors I will also add into my simulation. &lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - starts at 1.2%, what is official number given by https://www.usinflationcalculator.com/inflation/historical-inflation-rates/.&lt;br /&gt;
&lt;br /&gt;
Interest rates (short-term, fed-fund)  - starts at 0.25%, taken from https://tradingeconomics.com/united-states/interest-rate&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank / FED''&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20364</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20364"/>
		<updated>2021-01-17T13:53:11Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds) from banks in exchange for new money. The thought behind quantitativr easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable or banks to lend money to people, as with low interest rates and low, but permanent inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after wuantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment. In the second round of simulation, when banks will lend 90% of new money, I will add inflation rate variable, as it is suggested, that we did not face high level inflation (or maybe hyper-inflation) because banks did not lend majority of money they got. If they would have lend, it would be expected, that inflation would be much higher, as market demand for goods and services would be much higher. When I simulate behavior of stock owner, I refer to those owners, who do not trade very often and speculate on the stock markets. I refer to those, who mostly have bigger share in companies and are somehow involved in companies and tend to hold their stocks in long term.&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - rate of inflation will be important during 2. simulation, when i will pressume that majority of new money will be lend to households and companies, thus causing high rate of inflation. If inflation exceeds 2% year growth FED would stop QE.&lt;br /&gt;
&lt;br /&gt;
Crisis - flag, if there is a crisis FED is conducting open market operations using QE. After crisis ends there is period of reviving the economy when QE is conducted in higher rate.&lt;br /&gt;
&lt;br /&gt;
Probability of crisis - Higher the probability, higher the chance of FED going into new rounds of QE in order to stabilize the economy. When there is very low probability and the conomy is thriving, FED would just keep their assets at steady rate, with small purchases of US Treasuries after current ones expire.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank''&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20362</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20362"/>
		<updated>2021-01-17T11:13:08Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds) from banks in exchange for new money. The thought behind quantitativr easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable or banks to lend money to people, as with low interest rates and low, but permanent inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after wuantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;br /&gt;
&lt;br /&gt;
For this simulation I will use US economy enviroment.&lt;br /&gt;
&lt;br /&gt;
'''Global variables'''&lt;br /&gt;
&lt;br /&gt;
Inflation - rate of inflation will be important during 2. simulation, when i will pressume that majority of new money will be lend to households and companies, thus causing high rate of inflation. If inflation exceeds 2% year growth FED would stop QE.&lt;br /&gt;
&lt;br /&gt;
Crisis - flag, if there is a crisis FED is conducting open market operations using QE. After crisis ends there is period of reviving the economy when QE is conducted in higher rate.&lt;br /&gt;
&lt;br /&gt;
Probability of crisis - Higher the probability, higher the chance of FED going into new rounds of QE in order to stabilize the economy. When there is very low probability and the conomy is thriving, FED would just keep their assets at steady rate, with small purchases of US Treasuries after current ones expire.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Agents'''&lt;br /&gt;
&lt;br /&gt;
''Central Bank''&lt;br /&gt;
&lt;br /&gt;
''Banks''&lt;br /&gt;
&lt;br /&gt;
''Richest Stock Owners''&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20361</id>
		<title>Widening the spread between rich and poor</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Widening_the_spread_between_rich_and_poor&amp;diff=20361"/>
		<updated>2021-01-17T09:55:45Z</updated>

		<summary type="html">&lt;p&gt;Sára: Created page with &amp;quot;=Introduction=  After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach wa...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Introduction=&lt;br /&gt;
&lt;br /&gt;
After global economic crisis in 2007-2009 traditional approaches for stimulating economy growth failed. As a consequence of this a new experimental approach was established, called ''Quantitative easing''. A lot of people think that it is just 'printing money' and putting it into economy. But this is wrong and very narrow minded. In fact, the central bank of particular country buys financial assets (mostly governement bonds) from banks in exchange for new money. The thought behind quantitativr easing is also pretty straightforward, however numbers say that it ceases to work. The first weakness of it is that the money is not put straight into economy (for example through government spending), but it is given to banks. Now when we are in low interest rates situation, it is not very rentable or banks to lend money to people, as with low interest rates and low, but permanent inflation, they in fact get less value that they initially borrowed. So they often go and buy bonds and stocks themselves and thus overheat their prizes. And new money don't get into economy and so the growth of GDP is not promoted in a way, that it was thought. As 40% of stock market is owned by 5% of richest people in the world, it creates unequal distribution of wealth among population. Rich people get richer and poor people get even more poor and promotion of GDP growth stagnates. For example in the UK after wuantitative easing of 375 billion pounds led to 1.2-2% growth of its domestic GDP, what is in numbers 23-28 bn pounds. So we can see how highly ineffective this is.&lt;br /&gt;
&lt;br /&gt;
=Model=&lt;br /&gt;
&lt;br /&gt;
==Variables==&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20360</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20360"/>
		<updated>2021-01-14T22:04:24Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
Let's imagine now our automatic vacuum cleaner. It can be in a state, where it can decide from more options, which action to take. After taking each of available actions it can find itself in some other state with some probability. Sum over probabilities of all possible states, that the vacuum cleaner can end up in has to be equal to 1. For example if the vacuum cleaner is next to a sofa it can decide to either avoid it or try to get under the sofa and clean the area there. If it decides to go under the sofa there is 20% that if will not fit there and get damaged, 70% probability that it will clean it successfully and 10% probability that it will clean the area but get stuck there. So if we sum probabilities of all possible states after taking particular action it has to be 1. In our case 0.2 (next state is damaged) + 0.7 (next state next to a sofa after cleaning the area under it) + 0.1 (are is cleaned but it is stuck under the sofa) is equal to 1.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain and often less important. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&lt;br /&gt;
Firstly let's establish a function for determining value of a given state, when we have some policy set (in other words, at some state I obey the rule of taking the same action everytime).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This function makes to sum over all possible states s' i can end up (with some transition probability T(s, a, s') in where we sum expected immediate reward R after taking the action a and value of the same policy at state s' after taking discount factor into consideration.&lt;br /&gt;
&lt;br /&gt;
So now I my task is to find the optimal policy, the policy which gives me maximum reward.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_{\text{opt}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Here I don't have explicitly given action I need to take at state s, I compute value for all possible actions and choose the highest one.&lt;br /&gt;
&lt;br /&gt;
After computing maximum values for all states I have policy, which maps from each state to an action which maximizes reward, I can write it like this:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\pi_{\text{opt}}(s) = argmax_a \Bigl\{ \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')] \Bigr\}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Policy is a function for mapping from each state s from all possible states to an action a from all possible actions. In other words it optimizes the best action that should be taken from each state that an agent can end up in, in order to maximize the return function &amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt; (to get as big as reward as possible when facing possible infinite horizon). Because ending in some state is often not deterministic we follow a random path when we follow the policy. When defining a policy i need to use expected utility because of the randomness.&lt;br /&gt;
&lt;br /&gt;
=== Reinforcement Learning ===&lt;br /&gt;
When we want to talk about reinforcement learning we need to establish a new formula. In this case we will be in so called 'chance node' where we will have state and action defined.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Q_\pi(s, a) = \sum_{\text{s'}} T_a(s, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Reinforcement uses MDP when transition probability or reward is unknown. For example we use a generator that in every run generates new random probabilities for s'.&lt;br /&gt;
&lt;br /&gt;
When we calculate value of some state given some policy we can write the formula for the value using chance nodes. For example let's consider we have stocks of a company the company started to let out their employees. So I am in a state 'company is letting out its employees', my action is to sell stocks. There is probability 60% chance that stock prize will go down, in this case my reward is 10 as I avoided loosing money. However there is also 50% chance that it won't cause prize to drop but to  raise (reward -15, i could have made more money) and 10% is the chance, that it will more less stay the same with little fluctuations (reward 0, as I won't loose or gain if I sell). We can write formula for value at state 'company is letting out its emmployees' using chance node this way:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \begin{cases} Q_\pi(s, a) = \sum_{\text{s1'}} T_{\text{a1}}(s, s1') [R(s, a, s1') + \gamma V_\pi(s1')] \\ Q_\pi(s, a) = \sum_{\text{s2'}} T_{\text{a2}}(s, s2') [R(s, a, s2') + \gamma V_\pi(s2')] \\ Q_\pi(s, a) = \sum_{\text{s3'}} T_{\text{a3}}(s, s3') [R(s, a, s3') + \gamma V_\pi(s3')] \end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Where a1, a2, a3 are actions I can take at the current state and s1',s2',s3' are next states i can end up in after I take particular action.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Value Iteration ===&lt;br /&gt;
Using this algorithm we calculate the value of expected reward at the state s given that we use certain policy, for example i decide to alway steer right when I am at a crossroad.&lt;br /&gt;
The approach is this: Start with arbitrary policy values and repeatedly apply recurrences to converge to true values. However complicated this definition sounds, it is quite straightforward.&lt;br /&gt;
&lt;br /&gt;
In first step we initialize values for each policy at given state to 0.&lt;br /&gt;
&lt;br /&gt;
Then i start a loop of t iterations. For each iteration i will update value of each state depending on the value from previous iteration.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi^{\text{(t)}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi^{\text{(t-1)}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We iterate repeatedly values for all states until V converges with left-hand side equal to right hand side. When we make iterations for the whole policy we usually iterate until difference between value at state s' hasn't changed much from state s.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The time complexity of this algorithm is &amp;lt;math&amp;gt;O(tSS')&amp;lt;/math&amp;gt; because i iterate t-times over S possible states and these states are S' possible next states. The reason why we don't take actions into account is that i'm calculating the policy for all states and for each state i have exactly one action I take every time I end up in this particular actions, in other word my policy is already set.&lt;br /&gt;
When calculating optimal policy the complexity of this algorithm is &amp;lt;math&amp;gt;O(tSAS')&amp;lt;/math&amp;gt; as we take all possible actions available at given state into consideration and choose one that provides us with the biggest reward.&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
&lt;br /&gt;
State: at the start of a new round / out of the game&lt;br /&gt;
&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
&lt;br /&gt;
Rewards: 4 dollars, 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability. However there is also 33% probability of getting 4 dollars, but after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;br /&gt;
&lt;br /&gt;
Let's now set some policy and calculate value of that policy. So for example our policy is to take action 'stay', when in state 'at the start of a new round'.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(end)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
we know that the expected reward of policy 'end' is 0, as after the end of the game we won't get any more money, so let's write it into our formula.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{2}{3} (V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;12 = 4 + \frac{2}{3} (12)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 12&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
As the formula for finding the value of state given some policy says, we are looking for number (reward) when reward at state t is equal to the reward at time t - 1.&lt;br /&gt;
&lt;br /&gt;
''Let's now widen our example a little bit and see how it will change''&lt;br /&gt;
&lt;br /&gt;
Now after rolling a dice and the result is 1 or 2 I am in state when I have to toss a coin and if it is head I get 20 dollars and the game is over otherwise i loose 8 dollars and the game is over.&lt;br /&gt;
Here are the updates I have done:&lt;br /&gt;
&lt;br /&gt;
States: start of a new round, end of the game, tossing a coin&lt;br /&gt;
&lt;br /&gt;
Actions: stay, quit&lt;br /&gt;
&lt;br /&gt;
Rewards: 4, 10, 20, -8 dollars&lt;br /&gt;
&lt;br /&gt;
When I decide to quit nothing changes. If I decide to stay I get 4 dollars and can with probability of 67% start a new round or with 33% probability I'm in the state when I got 1 or 2 after rolling the dice and I toss a coin and with probability of 50% I will loose 8 dollars and end the game and same probability of gaining 12 dollars and ending the game.&lt;br /&gt;
Let's now calculate the expected value of state 'at the end of a new round' with the action 'stay'.&lt;br /&gt;
&lt;br /&gt;
If we still prefer policy which tells us to always choose action 'stay' when in state 'at the end of a new round', the value of this state given this policy changes&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(toss)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Expected value of toss is 0.5*-8 + 0.5*20= 6 and then the game ends, so it means that we won't have any further states that have to be taken into consideration.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{1}{3} (6) + \frac{2}{3} (V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 6 + \frac{2}{3} (V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;18 = 4 + \frac{1}{3} (6) + \frac{2}{3} (18)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 18&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Reward for the state 'at the end of a new round' given policy that we choose action 'stay' is 18.&lt;br /&gt;
&lt;br /&gt;
=Zdroje=&lt;br /&gt;
&lt;br /&gt;
==Articles==&lt;br /&gt;
[https://towardsdatascience.com/getting-started-with-markov-decision-processes-reinforcement-learning-ada7b4572ffb]&lt;br /&gt;
[https://www.worldscientific.com/doi/suppl/10.1142/p809/suppl_file/p809_chap01.pdf]&lt;br /&gt;
[https://towardsdatascience.com/reinforcement-learning-demystified-markov-decision-processes-part-1-bf00dda41690]&lt;br /&gt;
&lt;br /&gt;
==Videos==&lt;br /&gt;
[https://www.youtube.com/watch?v=9g32v7bK3Co&amp;amp;list=PLoROMvodv4rO1NB9TD4iUZ3qghGEGtqNX&amp;amp;index=7]&lt;br /&gt;
[https://www.youtube.com/watch?v=sdp49vTanSk]&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20359</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20359"/>
		<updated>2021-01-14T22:01:50Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
Let's imagine now our automatic vacuum cleaner. It can be in a state, where it can decide from more options, which action to take. After taking each of available actions it can find itself in some other state with some probability. Sum over probabilities of all possible states, that the vacuum cleaner can end up in has to be equal to 1. For example if the vacuum cleaner is next to a sofa it can decide to either avoid it or try to get under the sofa and clean the area there. If it decides to go under the sofa there is 20% that if will not fit there and get damaged, 70% probability that it will clean it successfully and 10% probability that it will clean the area but get stuck there. So if we sum probabilities of all possible states after taking particular action it has to be 1. In our case 0.2 (next state is damaged) + 0.7 (next state next to a sofa after cleaning the area under it) + 0.1 (are is cleaned but it is stuck under the sofa) is equal to 1.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain and often less important. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&lt;br /&gt;
Firstly let's establish a function for determining value of a given state, when we have some policy set (in other words, at some state I obey the rule of taking the same action everytime).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This function makes to sum over all possible states s' i can end up (with some transition probability T(s, a, s') in where we sum expected immediate reward R after taking the action a and value of the same policy at state s' after taking discount factor into consideration.&lt;br /&gt;
&lt;br /&gt;
So now I my task is to find the optimal policy, the policy which gives me maximum reward.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_{\text{opt}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Here I don't have explicitly given action I need to take at state s, I compute value for all possible actions and choose the highest one.&lt;br /&gt;
&lt;br /&gt;
After computing maximum values for all states I have policy, which maps from each state to an action which maximizes reward, I can write it like this:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\pi_{\text{opt}}(s) = argmax_a \Bigl\{ \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')] \Bigr\}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Policy is a function for mapping from each state s from all possible states to an action a from all possible actions. In other words it optimizes the best action that should be taken from each state that an agent can end up in, in order to maximize the return function &amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt; (to get as big as reward as possible when facing possible infinite horizon). Because ending in some state is often not deterministic we follow a random path when we follow the policy. When defining a policy i need to use expected utility because of the randomness.&lt;br /&gt;
&lt;br /&gt;
=== Reinforcement Learning ===&lt;br /&gt;
When we want to talk about reinforcement learning we need to establish a new formula. In this case we will be in so called 'chance node' where we will have state and action defined.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Q_\pi(s, a) = \sum_{\text{s'}} T_a(s, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Reinforcement uses MDP when transition probability or reward is unknown. For example we use a generator that in every run generates new random probabilities for s'.&lt;br /&gt;
&lt;br /&gt;
When we calculate value of some state given some policy we can write the formula for the value using chance nodes. For example let's consider we have stocks of a company the company started to let out their employees. So I am in a state 'company is letting out its employees', my action is to sell stocks. There is probability 60% chance that stock prize will go down, in this case my reward is 10 as I avoided loosing money. However there is also 50% chance that it won't cause prize to drop but to  raise (reward -15, i could have made more money) and 10% is the chance, that it will more less stay the same with little fluctuations (reward 0, as I won't loose or gain if I sell). We can write formula for value at state 'company is letting out its emmployees' using chance node this way:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \begin{cases} Q_\pi(s, a) = \sum_{\text{s1'}} T_{\text{a1}}(s, s1') [R(s, a, s1') + \gamma V_\pi(s1')] \\ Q_\pi(s, a) = \sum_{\text{s2'}} T_{\text{a2}}(s, s2') [R(s, a, s2') + \gamma V_\pi(s2')] \\ Q_\pi(s, a) = \sum_{\text{s3'}} T_{\text{a3}}(s, s3') [R(s, a, s3') + \gamma V_\pi(s3')] \end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Where a1, a2, a3 are actions I can take at the current state and s1',s2',s3' are next states i can end up in after I take particular action.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Value Iteration ===&lt;br /&gt;
Using this algorithm we calculate the value of expected reward at the state s given that we use certain policy, for example i decide to alway steer right when I am at a crossroad.&lt;br /&gt;
The approach is this: Start with arbitrary policy values and repeatedly apply recurrences to converge to true values. However complicated this definition sounds, it is quite straightforward.&lt;br /&gt;
&lt;br /&gt;
In first step we initialize values for each policy at given state to 0.&lt;br /&gt;
&lt;br /&gt;
Then i start a loop of t iterations. For each iteration i will update value of each state depending on the value from previous iteration.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi^{\text{(t)}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi^{\text{(t-1)}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We iterate repeatedly values for all states until V converges with left-hand side equal to right hand side. When we make iterations for the whole policy we usually iterate until difference between value at state s' hasn't changed much from state s.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The time complexity of this algorithm is &amp;lt;math&amp;gt;O(tSS')&amp;lt;/math&amp;gt; because i iterate t-times over S possible states and these states are S' possible next states. The reason why we don't take actions into account is that i'm calculating the policy for all states and for each state i have exactly one action I take every time I end up in this particular actions, in other word my policy is already set.&lt;br /&gt;
When calculating optimal policy the complexity of this algorithm is &amp;lt;math&amp;gt;O(tSAS')&amp;lt;/math&amp;gt; as we take all possible actions available at given state into consideration and choose one that provides us with the biggest reward.&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
&lt;br /&gt;
State: at the start of a new round / out of the game&lt;br /&gt;
&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
&lt;br /&gt;
Rewards: 4 dollars, 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability. However there is also 33% probability of getting 4 dollars, but after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;br /&gt;
&lt;br /&gt;
Let's now set some policy and calculate value of that policy. So for example our policy is to take action 'stay', when in state 'at the start of a new round'.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(end)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
we know that the expected reward of policy 'end' is 0, as after the end of the game we won't get any more money, so let's write it into our formula.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{2}{3} (V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;12 = 4 + \frac{2}{3} (12)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 12&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
As the formula for finding the value of state given some policy says, we are looking for number (reward) when reward at state t is equal to the reward at time t - 1.&lt;br /&gt;
&lt;br /&gt;
''Let's now widen our example a little bit and see how it will change''&lt;br /&gt;
&lt;br /&gt;
Now after rolling a dice and the result is 1 or 2 I am in state when I have to toss a coin and if it is head I get 20 dollars and the game is over otherwise i loose 8 dollars and the game is over.&lt;br /&gt;
Here are the updates I have done:&lt;br /&gt;
&lt;br /&gt;
States: start of a new round, end of the game, tossing a coin&lt;br /&gt;
Actions: stay, quit&lt;br /&gt;
Rewards: 4, 10, 20, -8 dollars&lt;br /&gt;
&lt;br /&gt;
When I decide to quit nothing changes. If I decide to stay I get 4 dollars and can with probability of 67% start a new round or with 33% probability I'm in the state when I got 1 or 2 after rolling the dice and I toss a coin and with probability of 50% I will loose 8 dollars and end the game and same probability of gaining 12 dollars and ending the game.&lt;br /&gt;
Let's now calculate the expected value of state 'at the end of a new round' with the action 'stay'.&lt;br /&gt;
&lt;br /&gt;
If we still prefer policy which tells us to always choose action 'stay' when in state 'at the end of a new round', the value of this state given this policy changes&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(toss)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Expected value of toss is 0.5*-8 + 0.5*20= 6 and then the game ends, so it means that we won't have any further states that have to be taken into consideration.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{1}{3} (6) + \frac{2}{3} (V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 6 + \frac{2}{3} (V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;18 = 4 + \frac{1}{3} (6) + \frac{2}{3} (18)&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 18&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Reward for the state 'at the end of a new round' given policy that we choose action 'stay' is 18.&lt;br /&gt;
&lt;br /&gt;
=Zdroje=&lt;br /&gt;
&lt;br /&gt;
==Articles==&lt;br /&gt;
[https://towardsdatascience.com/getting-started-with-markov-decision-processes-reinforcement-learning-ada7b4572ffb]&lt;br /&gt;
[https://www.worldscientific.com/doi/suppl/10.1142/p809/suppl_file/p809_chap01.pdf]&lt;br /&gt;
[https://towardsdatascience.com/reinforcement-learning-demystified-markov-decision-processes-part-1-bf00dda41690]&lt;br /&gt;
&lt;br /&gt;
==Videos==&lt;br /&gt;
[https://www.youtube.com/watch?v=9g32v7bK3Co&amp;amp;list=PLoROMvodv4rO1NB9TD4iUZ3qghGEGtqNX&amp;amp;index=7]&lt;br /&gt;
[https://www.youtube.com/watch?v=sdp49vTanSk]&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20358</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20358"/>
		<updated>2021-01-13T20:54:31Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
Let's imagine now our automatic vacuum cleaner. It can be in a state, where it can decide from more options, which action to take. After taking each of available actions it can find itself in some other state with some probability. Sum over probabilities of all possible states, that the vacuum cleaner can end up in has to be equal to 1. For example if the vacuum cleaner is next to a sofa it can decide to either avoid it or try to get under the sofa and clean the area there. If it decides to go under the sofa there is 20% that if will not fit there and get damaged, 70% probability that it will clean it successfully and 10% probability that it will clean the area but get stuck there. So if we sum probabilities of all possible states after taking particular action it has to be 1. In our case 0.2 (next state is damaged) + 0.7 (next state next to a sofa after cleaning the area under it) + 0.1 (are is cleaned but it is stuck under the sofa) is equal to 1.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain it is. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&amp;lt;math&amp;gt;\pi (a | s) = P[A_t = a | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Policy is a function for mapping from each state s from all possible states to an action a from all possible actions. In other words it optimizes the best action that should be taken from each state that an agent can end up in, in order to maximize the return function &amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt; (to get as big as reward as possible when facing possible infinite horizon). Because ending in some state is often not deterministic we follow a random path when we follow the policy. When defining a policy i need to use expected utility because of the randomness.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This function tells me the value of policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; applied at state s, which is equal to sum over all possible states s' i can end up (with some transition probability T(s, a, s') in where we sum expected immediate reward R after taking the action a and value of the same policy at state s' after taking discount factor into consideration.&lt;br /&gt;
&lt;br /&gt;
So now I my task is to find the optimal policy, the policy which gives me maximum reward.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_{\text{opt}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
For this maximum reward function I get optimum policy function as follows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\pi_{\text{opt}}(s) = argmax_a \Bigl\{ \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')] \Bigr\}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Reinforcement Learning ===&lt;br /&gt;
When we want to talk about reinforcement learning we need to establish a new formula. In this case we will be in so called 'chance node' where we will have state and action defined.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Q_\pi(s, a) = \sum_{\text{s'}} T_a(s, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Reinforcement uses MDP when transition probability or reward is unknown. For example we use a generator that in every run generates new random probabilities for s'.&lt;br /&gt;
&lt;br /&gt;
When we calculate value of some state given some policy we can write the formula for the value using chance nodes. For example let's consider we have stocks of a company the company started to let out their employees. So I am in a state 'company is letting out its employees', my action is to sell stocks. There is probability 60% chance that stock prize will go down, in this case my reward is 10 as I avoided loosing money. However there is also 50% chance that it won't cause prize to drop but to  raise (reward -15, i could have made more money) and 10% is the chance, that it will more less stay the same with little fluctuations (reward 0, as I won't loose or gain if I sell). We can write formula for value at state 'company is letting out its emmployees' using chance node this way:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \begin{cases} Q_\pi(s, a) = \sum_{\text{s1'}} T_{\text{a1}}(s, s1') [R(s, a, s1') + \gamma V_\pi(s1')] \\ Q_\pi(s, a) = \sum_{\text{s2'}} T_{\text{a2}}(s, s2') [R(s, a, s2') + \gamma V_\pi(s2')] \\ Q_\pi(s, a) = \sum_{\text{s3'}} T_{\text{a3}}(s, s3') [R(s, a, s3') + \gamma V_\pi(s3')] \end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Where a1, a2, a3 are actions I can take at the current state and s1',s2',s3' are next states i can end up in after I take particular action.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Value Iteration ===&lt;br /&gt;
Using this algorithm we calculate the value of expected reward at the state s given that we use certain policy, for example i decide to alway steer right when I am at a crossroad.&lt;br /&gt;
The approach is this: Start with arbitrary policy values and repeatedly apply recurrences to converge to true values. However complicated this definition sounds, it is quite straightforward.&lt;br /&gt;
&lt;br /&gt;
In first step we initialize values for each policy at given state to 0.&lt;br /&gt;
&lt;br /&gt;
Then i start a loop of t iterations. For each iteration i will update value of each state depending on the value from previous iteration.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi^{\text{(t)}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi^{\text{(t-1)}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We iterate repeatedly values for all states until V converges with left-hand side equal to right hand side. When we make iterations for the whole policy we usually iterate until difference between value at state s' hasn't changed much from state s.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The time complexity of this algorithm is &amp;lt;math&amp;gt;O(tSS')&amp;lt;/math&amp;gt; because i iterate t-times over S possible states and these states are S' possible next states. The reason why we don't take actions into account is that i'm calculating the policy for all states and for each state i have exactly one action I take every time I end up in this particular actions, in other word my policy is already set.&lt;br /&gt;
When calculating optimal policy the complexity of this algorithm is &amp;lt;math&amp;gt;O(tSAS')&amp;lt;/math&amp;gt; as we take all possible actions available at given state into consideration and choose one that provides us with the biggest reward.&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
&lt;br /&gt;
State: at the start of a new round / out of the game&lt;br /&gt;
&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
&lt;br /&gt;
Rewards: 4 dollars, 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability. However there is also 33% probability of getting 4 dollars, but after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;br /&gt;
&lt;br /&gt;
Let's now set some policy and calculate value of that policy. So for example our policy is to take action 'stay', when in state 'at the start of a new round'.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(end)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
we know that the expected reward of policy 'end' is 0, as after the end of the game we won't get any more money, so let's write it into our formula.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;4 = \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 12&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
''Let's now widen our example a little bit and see how it will change''&lt;br /&gt;
&lt;br /&gt;
Now after rolling a dice and the result is 1 or 2 I am in state when I have to toss a coin and if it is head I get 12 dollars and the game is over otherwise i loose 8 dollars and the game is over.&lt;br /&gt;
Here are the updates I have done:&lt;br /&gt;
&lt;br /&gt;
States: start of a new round, end of the game, tossing a coin&lt;br /&gt;
Actions: stay, quit&lt;br /&gt;
Rewards: 4, 10, 12, -8 dollars&lt;br /&gt;
&lt;br /&gt;
When I decide to quit nothing changes. If I decide to stay I get 4 dollars and can with probability of 67% start a new round or with 33% probability I'm in the state when I got 1 or 2 after rolling the dice and I toss a coin and with probability of 50% I will loose 8 dollars and end the game and same probability of gaining 12 dollars and ending the game.&lt;br /&gt;
Let's now calculate the expected value of state 'at the end of a new round' with the action 'stay'.&lt;br /&gt;
&lt;br /&gt;
If we still prefer policy which tells us to always choose action 'stay' when in state 'at the end of a new round', the value of this state given this policy changes&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(toss)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Expected value of toss is 0.5*-8 + 0.5*12= 2.&lt;br /&gt;
&lt;br /&gt;
=Zdroje=&lt;br /&gt;
&lt;br /&gt;
==Articles==&lt;br /&gt;
[https://towardsdatascience.com/getting-started-with-markov-decision-processes-reinforcement-learning-ada7b4572ffb]&lt;br /&gt;
[https://www.worldscientific.com/doi/suppl/10.1142/p809/suppl_file/p809_chap01.pdf]&lt;br /&gt;
[https://towardsdatascience.com/reinforcement-learning-demystified-markov-decision-processes-part-1-bf00dda41690]&lt;br /&gt;
&lt;br /&gt;
==Videos==&lt;br /&gt;
[https://www.youtube.com/watch?v=9g32v7bK3Co&amp;amp;list=PLoROMvodv4rO1NB9TD4iUZ3qghGEGtqNX&amp;amp;index=7]&lt;br /&gt;
[https://www.youtube.com/watch?v=sdp49vTanSk]&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20357</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20357"/>
		<updated>2021-01-12T22:12:45Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
Let's imagine now our automatic vacuum cleaner. It can be in a state, where it can decide from more options, which action to take. After taking each of available actions it can find itself in some other state with some probability. Sum over probabilities of all possible states, that the vacuum cleaner can end up in has to be equal to 1. For example if the vacuum cleaner is next to a sofa it can decide to either avoid it or try to get under the sofa and clean the area there. If it decides to go under the sofa there is 20% that if will not fit there and get damaged, 70% probability that it will clean it successfully and 10% probability that it will clean the area but get stuck there. So if we sum probabilities of all possible states after taking particular action it has to be 1. In our case 0.2 (next state is damaged) + 0.7 (next state next to a sofa after cleaning the area under it) + 0.1 (are is cleaned but it is stuck under the sofa) is equal to 1.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain it is. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&amp;lt;math&amp;gt;\pi (a | s) = P[A_t = a | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Policy is a function for mapping from each state s from all possible states to an action a from all possible actions. In other words it optimizes the best action that should be taken from each state that an agent can end up in, in order to maximize the return function &amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt; (to get as big as reward as possible when facing possible infinite horizon). Because ending in some state is often not deterministic we follow a random path when we follow the policy. When defining a policy i need to use expected utility because of the randomness.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This function tells me the value of policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; applied at state s, which is equal to sum over all possible states s' i can end up (with some transition probability T(s, a, s') in where we sum expected immediate reward R after taking the action a and value of the same policy at state s' after taking discount factor into consideration.&lt;br /&gt;
&lt;br /&gt;
So now I my task is to find the optimal policy, the policy which gives me maximum reward.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_{\text{opt}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
For this maximum reward function I get optimum policy function as follows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\pi_{\text{opt}}(s) = argmax_a \Bigl\{ \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')] \Bigr\}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Reinforcement Learning ===&lt;br /&gt;
When we want to talk about reinforcement learning we need to establish a new formula. In this case we will be in so called 'chance node' where we will have state and action defined.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Q_\pi(s, a) = \sum_{\text{s'}} T_a(s, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Reinforcement uses MDP when transition probability or reward is unknown. For example we use a generator that in every run generates new random probabilities for s'.&lt;br /&gt;
&lt;br /&gt;
When we calculate value of some state given some policy we can write the formula for the value using chance nodes. For example let's consider we have stocks of a company the company started to let out their employees. So I am in a state 'company is letting out its employees', my action is to sell stocks. There is probability 60% chance that stock prize will go down, in this case my reward is 10 as I avoided loosing money. However there is also 50% chance that it won't cause prize to drop but to  raise (reward -15, i could have made more money) and 10% is the chance, that it will more less stay the same with little fluctuations (reward 0, as I won't loose or gain if I sell). We can write formula for value at state 'company is letting out its emmployees' using chance node this way:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \begin{cases} Q_\pi(s, a) = \sum_{\text{s1'}} T_{\text{a1}}(s, s1') [R(s, a, s1') + \gamma V_\pi(s1')] \\ Q_\pi(s, a) = \sum_{\text{s2'}} T_{\text{a2}}(s, s2') [R(s, a, s2') + \gamma V_\pi(s2')] \\ Q_\pi(s, a) = \sum_{\text{s3'}} T_{\text{a3}}(s, s3') [R(s, a, s3') + \gamma V_\pi(s3')] \end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Where a1, a2, a3 are actions I can take at the current state and s1',s2',s3' are next states i can end up in after I take particular action.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Value Iteration ===&lt;br /&gt;
Using this algorithm we calculate the value of expected reward at the state s given that we use certain policy, for example i decide to alway steer right when I am at a crossroad.&lt;br /&gt;
The approach is this: Start with arbitrary policy values and repeatedly apply recurrences to converge to true values. However complicated this definition sounds, it is quite straightforward.&lt;br /&gt;
&lt;br /&gt;
In first step we initialize values for each policy at given state to 0.&lt;br /&gt;
&lt;br /&gt;
Then i start a loop of t iterations. For each iteration i will update value of each state depending on the value from previous iteration.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi^{\text{(t)}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi^{\text{(t-1)}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We iterate repeatedly values for all states until V converges with left-hand side equal to right hand side. When we make iterations for the whole policy we usually iterate until difference between value at state s' hasn't changed much from state s.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The time complexity of this algorithm is &amp;lt;math&amp;gt;O(tSS')&amp;lt;/math&amp;gt; because i iterate t-times over S possible states and these states are S' possible next states. The reason why we don't take actions into account is that i'm calculating the policy for all states and for each state i have exactly one action I take every time I end up in this particular actions, in other word my policy is already set.&lt;br /&gt;
When calculating optimal policy the complexity of this algorithm is &amp;lt;math&amp;gt;O(tSAS')&amp;lt;/math&amp;gt; as we take all possible actions available at given state into consideration and choose one that provides us with the biggest reward.&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
&lt;br /&gt;
State: at the start of a new round / out of the game&lt;br /&gt;
&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
&lt;br /&gt;
Rewards: 4 dollars, 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability. However there is also 33% probability of getting 4 dollars, but after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;br /&gt;
&lt;br /&gt;
Let's now set some policy and calculate value of that policy. So for example our policy is to take action 'stay', when in state 'at the start of a new round'.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(end)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
we know that the expected reward of policy 'stay' is 0, as after the end of the game we won't get any more money, so let's write it into our formula.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;4 = \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 12&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
''Let's now widen our example a little bit and see how it will change''&lt;br /&gt;
&lt;br /&gt;
Now after rolling a dice and the result is 1 or 2 I am in state when I have to toss a coin and if it is head I get 15 dollars and the game is over otherwise i loose 8 dollars and the game is over or I can ask to roll dice again but loosing 1 dollar.&lt;br /&gt;
Here are the updates I have done:&lt;br /&gt;
&lt;br /&gt;
States: start of a new round, end of the game, dice resulted in 1 or 2&lt;br /&gt;
Actions: stay, quit, toss a coin, ask to roll the dice again&lt;br /&gt;
Rewards: 4, 10, 15, -8, -1 dollars&lt;br /&gt;
&lt;br /&gt;
When I decide to quit nothing changes. If I decide to stay I get 4 dollars and can with probability of 67% start a new round or with 33% probability I'm in the state when I got 1 or 2 after rolling the dice and I can either roll the dice again and have the same probabilities and same states as in case of stay action but I will loose 1 dollar, or I can try to toss a coin and with probability of 50% I will loose 8 dollars and end the game and same probability of gaining 15 dollars and ending the game.&lt;br /&gt;
Let's now calculate the expected value of state 'at the end of a new round' with the action 'stay'.&lt;br /&gt;
&lt;br /&gt;
=Zdroje=&lt;br /&gt;
&lt;br /&gt;
==Articles==&lt;br /&gt;
[https://towardsdatascience.com/getting-started-with-markov-decision-processes-reinforcement-learning-ada7b4572ffb]&lt;br /&gt;
[https://www.worldscientific.com/doi/suppl/10.1142/p809/suppl_file/p809_chap01.pdf]&lt;br /&gt;
[https://towardsdatascience.com/reinforcement-learning-demystified-markov-decision-processes-part-1-bf00dda41690]&lt;br /&gt;
&lt;br /&gt;
==Videos==&lt;br /&gt;
[https://www.youtube.com/watch?v=9g32v7bK3Co&amp;amp;list=PLoROMvodv4rO1NB9TD4iUZ3qghGEGtqNX&amp;amp;index=7]&lt;br /&gt;
[https://www.youtube.com/watch?v=sdp49vTanSk]&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20330</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20330"/>
		<updated>2021-01-09T22:32:44Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
Let's imagine now our automatic vacuum cleaner. It can be in a state, where it can decide from more options, which action to take. After taking each of available actions it can find itself in some other state with some probability. Sum over probabilities of all possible states, that the vacuum cleaner can end up in has to be equal to 1. For example if the vacuum cleaner is next to a sofa it can decide to either avoid it or try to get under the sofa and clean the area there. If it decides to go under the sofa there is 20% that if will not fit there and get damaged, 70% probability that it will clean it successfully and 10% probability that it will clean the area but get stuck there. So if we sum probabilities of all possible states after taking particular action it has to be 1. In our case 0.2 (next state is damaged) + 0.7 (next state next to a sofa after cleaning the area under it) + 0.1 (are is cleaned but it is stuck under the sofa) is equal to 1.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain it is. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&amp;lt;math&amp;gt;\pi (a | s) = P[A_t = a | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Policy is a function for mapping from each state s from all possible states to an action a from all possible actions. In other words it optimizes the best action that should be taken from each state that an agent can end up in, in order to maximize the return function &amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt; (to get as big as reward as possible when facing possible infinite horizon). Because ending in some state is often not deterministic we follow a random path when we follow the policy. When defining a policy i need to use expected utility because of the randomness.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This function tells me the value of policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; applied at state s, which is equal to sum over all possible states s' i can end up (with some transition probability T(s, a, s') in where we sum expected immediate reward R after taking the action a and value of the same policy at state s' after taking discount factor into consideration.&lt;br /&gt;
&lt;br /&gt;
So now I my task is to find the optimal policy, the policy which gives me maximum reward.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_{\text{opt}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
For this maximum reward function I get optimum policy function as follows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\pi_{\text{opt}}(s) = argmax_a \Bigl\{ \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')] \Bigr\}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Reinforcement Learning ===&lt;br /&gt;
When we want to talk about reinforcement learning we need to establish a new formula. In this case we will be in so called 'chance node' where we will have state and action defined.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Q_\pi(s, a) = \sum_{\text{s'}} T_a(s, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Reinforcement uses MDP when transition probability or reward is unknown. For example we use a generator that in every run generates new random probabilities for s'.&lt;br /&gt;
&lt;br /&gt;
When we calculate value of some state given some policy we can write the formula for the value using chance nodes. For example let's consider we have stocks of a company the company started to let out their employees. So I am in a state 'company is letting out its employees', my action is to sell stocks. There is probability 60% chance that stock prize will go down, in this case my reward is 10 as I avoided loosing money. However there is also 50% chance that it won't cause prize to drop but to  raise (reward -15, i could have made more money) and 10% is the chance, that it will more less stay the same with little fluctuations (reward 0, as I won't loose or gain if I sell). We can write formula for value at state 'company is letting out its emmployees' using chance node this way:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \begin{cases} Q_\pi(s, a) = \sum_{\text{s1'}} T_{\text{a1}}(s, s1') [R(s, a, s1') + \gamma V_\pi(s1')] \\ Q_\pi(s, a) = \sum_{\text{s2'}} T_{\text{a2}}(s, s2') [R(s, a, s2') + \gamma V_\pi(s2')] \\ Q_\pi(s, a) = \sum_{\text{s3'}} T_{\text{a3}}(s, s3') [R(s, a, s3') + \gamma V_\pi(s3')] \end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Where a1, a2, a3 are actions I can take at the current state and s1',s2',s3' are next states i can end up in after I take particular action.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Value Iteration ===&lt;br /&gt;
Using this algorithm we calculate the value of expected reward at the state s given that we use certain policy, for example i decide to alway steer right when I am at a crossroad.&lt;br /&gt;
The approach is this: Start with arbitrary policy values and repeatedly apply recurrences to converge to true values. However complicated this definition sounds, it is quite straightforward.&lt;br /&gt;
&lt;br /&gt;
In first step we initialize values for each policy at given state to 0.&lt;br /&gt;
&lt;br /&gt;
Then i start a loop of t iterations. For each iteration i will update value of each state depending on the value from previous iteration.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi^{\text{(t)}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi^{\text{(t-1)}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We iterate repeatedly values for all states until V converges with left-hand side equal to right hand side. When we make iterations for the whole policy we usually iterate until difference between value at state s' hasn't changed much from state s.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The time complexity of this algorithm is &amp;lt;math&amp;gt;O(tSS')&amp;lt;/math&amp;gt; because i iterate t-times over S possible states and these states are S' possible next states. The reason why we don't take actions into account is that i'm calculating the policy for all states and for each state i have exactly one action I take every time I end up in this particular actions, in other word my policy is already set.&lt;br /&gt;
When calculating optimal policy the complexity of this algorithm is &amp;lt;math&amp;gt;O(tSAS')&amp;lt;/math&amp;gt; as we take all possible actions available at given state into consideration and choose one that provides us with the biggest reward.&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
&lt;br /&gt;
State: at the start of a new round / out of the game&lt;br /&gt;
&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
&lt;br /&gt;
Rewards: 4 dollars, 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability. However there is also 33% probability of getting 4 dollars, but after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;br /&gt;
&lt;br /&gt;
Let's now set some policy and calculate value of that policy. So for example our policy is to take action 'stay', when in state 'at the start of a new round'.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(end)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
we know that the expected reward of policy 'stay' is 0, as after the end of the game we won't get any more money, so let's write it into our formula.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;4 = \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 12&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
''Let's now widen our example a little bit and see how it will change''&lt;br /&gt;
&lt;br /&gt;
Now after rolling a dice and the result is 1 or 2 I am in state when I have to toss a coin and if it is head I get 15 dollars and the game is over otherwise i loose 8 dollars and the game is over or I can ask to roll dice again but loosing 1 dollar.&lt;br /&gt;
Here are the updates I have done:&lt;br /&gt;
&lt;br /&gt;
States: start of a new round, end of the game, dice resulted in 1 or 2&lt;br /&gt;
Actions: stay, quit, toss a coin, ask to roll the dice again&lt;br /&gt;
Rewards: 4, 10, 15, -8, -1 dollars&lt;br /&gt;
&lt;br /&gt;
When I decide to quit nothing changes. If I decide to stay I get 4 dollars and can with probability of 67% start a new round or with 33% probability I'm in the state when I got 1 or 2 after rolling the dice and I can either roll the dice again and have the same probabilities and same states as in case of stay action but I will loose 1 dollar, or I can try to toss a coin and with probability of 50% I will loose 8 dollars and end the game and same probability of gaining 15 dollars and ending the game.&lt;br /&gt;
Let's now calculate the expected value of state 'at the end of a new round' with the action 'stay'.&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20329</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20329"/>
		<updated>2021-01-09T22:23:35Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
Let's imagine now our automatic vacuum cleaner. It can be in a state, where it can decide from more options, which action to take. After taking each of available actions it can find itself in some other state with some probability. Sum over probabilities of all possible states, that the vacuum cleaner can end up in has to be equal to 1. For example if the vacuum cleaner is next to a sofa it can decide to either avoid it or try to get under the sofa and clean the area there. If it decides to go under the sofa there is 20% that if will not fit there and get damaged, 70% probability that it will clean it successfully and 10% probability that it will clean the area but get stuck there. So if we sum probabilities of all possible states after taking particular action it has to be 1. In our case 0.2 (next state is damaged) + 0.7 (next state next to a sofa after cleaning the area under it) + 0.1 (are is cleaned but it is stuck under the sofa) is equal to 1.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain it is. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&amp;lt;math&amp;gt;\pi (a | s) = P[A_t = a | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Policy is a function for mapping from each state s from all possible states to an action a from all possible actions. In other words it optimizes the best action that should be taken from each state that an agent can end up in, in order to maximize the return function &amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt; (to get as big as reward as possible when facing possible infinite horizon). Because ending in some state is often not deterministic we follow a random path when we follow the policy. When defining a policy i need to use expected utility because of the randomness.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This function tells me the value of policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; applied at state s, which is equal to sum over all possible states s' i can end up (with some transition probability T(s, a, s') in where we sum expected immediate reward R after taking the action a and value of the same policy at state s' after taking discount factor into consideration.&lt;br /&gt;
&lt;br /&gt;
So now I my task is to find the optimal policy, the policy which gives me maximum reward.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_{\text{opt}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
For this maximum reward function I get optimum policy function as follows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\pi_{\text{opt}}(s) = argmax_a \Bigl\{ \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')] \Bigr\}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Reinforcement Learning ===&lt;br /&gt;
When we want to talk about reinforcement learning we need to establish a new formula. In this case we will be in so called 'chance node' where we will have state and action defined.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;Q_\pi(s, a) = \sum_{\text{s'}} T_a(s, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Reinforcement uses MDP when transition probability or reward is unknown. For example we use a generator that in every run generates new random probabilities for s'.&lt;br /&gt;
&lt;br /&gt;
When we calculate value of some state given some policy we can write the formula for the value using chance nodes. For example let's consider we have stocks of a company the company started to let out their employees. So I am in a state 'company is letting out its employees', my action is to sell stocks. There is probability 60% chance that stock prize will go down, in this case my reward is 10 as I avoided loosing money. However there is also 50% chance that it won't cause prize to drop but to  raise (reward -15, i could have made more money) and 10% is the chance, that it will more less stay the same with little fluctuations (reward 0, as I won't loose or gain if I sell). We can write formula for value at state 'company is letting out its emmployees' using chance node this way:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \begin{cases} Q_\pi(s, a) = \sum_{\text{s1'}} T_{\text{a1}}(s, s1') [R(s, a, s1') + \gamma V_\pi(s1')] \\ Q_\pi(s, a) = \sum_{\text{s2'}} T_{\text{a2}}(s, s2') [R(s, a, s2') + \gamma V_\pi(s2')] \\ Q_\pi(s, a) = \sum_{\text{s3'}} T_{\text{a3}}(s, s3') [R(s, a, s3') + \gamma V_\pi(s3')] \end{cases}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Where a1, a2, a3 are actions I can take at the current state and s1',s2',s3' are next states i can end up in after I take particular action.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Value Iteration ===&lt;br /&gt;
Using this algorithm we calculate the value of expected reward at the state s given that we use certain policy, for example i decide to alway steer right when I am at a crossroad.&lt;br /&gt;
The approach is this: Start with arbitrary policy values and repeatedly apply recurrences to converge to true values. However complicated this definition sounds, it is quite straightforward.&lt;br /&gt;
&lt;br /&gt;
In first step we initialize values for each policy at given state to 0.&lt;br /&gt;
&lt;br /&gt;
Then i start a loop of t iterations. For each iteration i will update value of each state depending on the value from previous iteration.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi^{\text{(t)}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi^{\text{(t-1)}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We iterate repeatedly values for all states until V converges with left-hand side equal to right hand side. When we make iterations for the whole policy we usually iterate until difference between value at state s' hasn't changed much from state s.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The time complexity of this algorithm is &amp;lt;math&amp;gt;O(tSS')&amp;lt;/math&amp;gt; because i iterate t-times over S possible states and these states are S' possible next states. The reason why we don't take actions into account is that i'm calculating the policy for all states and for each state i have exactly one action I take every time I end up in this particular actions, in other word my policy is already set.&lt;br /&gt;
When calculating optimal policy the complexity of this algorithm is &amp;lt;math&amp;gt;O(tSAS')&amp;lt;/math&amp;gt; as we take all possible actions available at given state into consideration and choose one that provides us with the biggest reward.&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
&lt;br /&gt;
State: at the start of a new round / out of the game&lt;br /&gt;
&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
&lt;br /&gt;
Reward: 4 dollars / 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability. However there is also 33% probability of getting 4 dollars, but after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;br /&gt;
&lt;br /&gt;
Let's now set some policy and calculate value of that policy. So for example our policy is to take action 'stay', when in state 'at the start of a new round'.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(end)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
we know that the expected reward of policy 'stay' is 0, as after the end of the game we won't get any more money, so let's write it into our formula.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;4 = \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 12&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
''Let's now widen our example a little bit and see how it will change''&lt;br /&gt;
&lt;br /&gt;
Now after rolling a dice and the result is 1 or 2 I am in state when I have to toss a coin and if it is head I get 15 dollars and the game is over otherwise i loose 8 dollars and the game is over.&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20328</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20328"/>
		<updated>2021-01-09T21:26:15Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
Let's imagine now our automatic vacuum cleaner. It can be in a state, where it can decide from more options, which action to take. After taking each of available actions it can find itself in some other state with some probability. Sum over probabilities of all possible states, that the vacuum cleaner can end up in has to be equal to 1. For example if the vacuum cleaner is next to a sofa it can decide to either avoid it or try to get under the sofa and clean the area there. If it decides to go under the sofa there is 20% that if will not fit there and get damaged, 70% probability that it will clean it successfully and 10% probability that it will clean the area but get stuck there. So if we sum probabilities of all possible states after taking particular action it has to be 1. In our case 0.2 (next state is damaged) + 0.7 (next state next to a sofa after cleaning the area under it) + 0.1 (are is cleaned but it is stuck under the sofa) is equal to 1.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain it is. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&amp;lt;math&amp;gt;\pi (a | s) = P[A_t = a | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Policy is a function for mapping from each state s from all possible states to an action a from all possible actions. In other words it optimizes the best action that should be taken from each state that an agent can end up in, in order to maximize the return function &amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt; (to get as big as reward as possible when facing possible infinite horizon). Because ending in some state is often not deterministic we follow a random path when we follow the policy. When defining a policy i need to use expected utility because of the randomness.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This function tells me the value of policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; applied at state s, which is equal to sum over all possible states s' i can end up (with some transition probability T(s, a, s') in where we sum expected immediate reward R after taking the action a and value of the same policy at state s' after taking discount factor into consideration.&lt;br /&gt;
&lt;br /&gt;
So now I my task is to find the optimal policy, the policy which gives me maximum reward.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_{\text{opt}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
For this maximum reward function I get optimum policy function as follows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\pi_{\text{opt}}(s) = argmax_a \Bigl\{ \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_{\text{opt}}(s')] \Bigr\}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Value Iteration ===&lt;br /&gt;
Using this algorithm we calculate the value of expected reward at the state s given that we use certain policy, for example i decide to alway steer right when I am at a crossroad.&lt;br /&gt;
The approach is this: Start with arbitrary policy values and repeatedly apply recurrences to converge to true values. However complicated this definition sounds, it is quite straightforward.&lt;br /&gt;
&lt;br /&gt;
In first step we initialize values for each policy at given state to 0.&lt;br /&gt;
&lt;br /&gt;
Then i start a loop of t iterations. For each iteration i will update value of each state depending on the value from previous iteration.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi^{\text{(t)}}(s) = \sum_{\text{s'}} T(s, a, s') [R(s, a, s') + \gamma V_\pi^{\text{(t-1)}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We iterate repeatedly values for all states until V converges with left-hand side equal to right hand side. When we make iterations for the whole policy we usually iterate until difference between value at state s' hasn't changed much from state s.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The time complexity of this algorithm is &amp;lt;math&amp;gt;O(tSS')&amp;lt;/math&amp;gt; because i iterate t-times over S possible states and these states are S' possible next states. The reason why we don't take actions into account is that i'm calculating the policy for all states and for each state i have exactly one action I take every time I end up in this particular actions.&lt;br /&gt;
When calculating optimal policy the complexity of this algorithm is &amp;lt;math&amp;gt;O(tSAS')&amp;lt;/math&amp;gt; as we take all possible actions available at given state into consideration and choose one that provides us with the biggest reward.&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example 1 ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
&lt;br /&gt;
State: at the start of a new round / out of the game&lt;br /&gt;
&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
&lt;br /&gt;
Reward: 4 dollars / 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability. However there is also 33% probability of getting 4 dollars, but after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;br /&gt;
&lt;br /&gt;
Let's now set some policy and calculate value of that policy. So for example our policy is to take action 'stay', when in state 'at the start of a new round'.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(end)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
we know that the expected reward of policy 'stay' is 0, as after the end of the game we won't get any more money, so let's write it into our formula.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;4 = \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 12&amp;lt;/math&amp;gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20327</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20327"/>
		<updated>2021-01-08T22:28:09Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
 Let's imagine now our automatic vacuum cleaner. It can be in a state, where it can decide from more options, which action to take. After taking each of available actions it can find itself in some other state with some probability. Sum over probabilities of all possible states, that the vacuum cleaner can end up in has to be equal to 1. For example if the vacuum cleaner is next to a sofa it can decide to either avoid it or try to get under the sofa and clean the area there. If it decides to go under the sofa there is 20% that if will not fit there and get damaged, 70% probability that it will clean it successfully and 10% probability that it will clean the area but get stuck there. So if we sum probabilities of all possible states after taking particular action it has to be 1. In our case 0.2 (next state is damaged) + 0.7 (next state next to a sofa after cleaning the area under it) + 0.1 (are is cleaned but it is stuck under the sofa) is equal to 1.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain it is. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&amp;lt;math&amp;gt;\pi (a | s) = P[A_t = a | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Policy is a function for mapping from each states s from all possible states to an action a from all possible actions. In other words it optimizes the best action that should be taken from each state that an agent can end up in, in order to maximize the return function mentioned above (to get as big as reward as possible when facing possible infinite horizon). Because ending in some state is often not deterministic we follow a random path when we follow the policy. When defining a policy i need to use expected utility because of the randomness.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \sum_s' T(s, a, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This function tells me the value of policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; applied at state s, which is equal to sum over all possible states s' i can end up in where we sum expected immediate reward R after taking the action a and value of the same policy at state s' after taking discount factor into consideration.&lt;br /&gt;
&lt;br /&gt;
=== Value Iteration ===&lt;br /&gt;
Using this algorithm we calculate the value of each policy we can apply at given state s.&lt;br /&gt;
The approach is this: Start with arbitrary policy values and repeatedly apply recurrences to converge to true values. However complicated this definition sounds, it is quite straightforward.&lt;br /&gt;
&lt;br /&gt;
In first step we initialize values for each policy at given state to 0.&lt;br /&gt;
&lt;br /&gt;
Then i start a loop of t iterations. For each iteration i will update value of each state depending on the value from previous iteration.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi^{\text{(t)}}(s) = \sum_s' T(s, a, s') [R(s, a, s') + \gamma V_\pi^{\text{(t-1)}}(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
When we consider the value of t, which stands for number of iterations. We mostly set the t with reflection to change in values of given policy during iterations. So if the change drops under some value i stop the iteration and i take the results from my last iteration as values of all policies given all possible states.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The time complexity of this algorithm is &amp;lt;math&amp;gt;O(tSS')&amp;lt;/math&amp;gt; because i iterate t-times over S possible states and these states are S' possible next states. The reason why we don't take actions into account is that i'm calculating the policy for all states and for each state i have exactly one action I take every time I end up in this particular actions.&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example 1 ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
&lt;br /&gt;
State: at the start of a new round / out of the game&lt;br /&gt;
&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
&lt;br /&gt;
Reward: 4 dollars / 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability. However there is also 33% probability of getting 4 dollars, but after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;br /&gt;
&lt;br /&gt;
Let's now set some policy and calculate value of that policy. So for example our policy is to take action 'stay', when in state 'at the start of a new round'.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(end)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
we know that the expected reward of policy 'stay' is 0, as after the end of the game we won't get any more money, so let's write it into our formula.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;4 = \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 12&amp;lt;/math&amp;gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20326</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20326"/>
		<updated>2021-01-08T21:58:34Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
 Let's imagine now our automatic vacuum cleaner. It can be in a state, where it can decide from more options, which action to take. After taking each of available actions it can find itself in some other state with some probability. Sum over probabilities of all possible states, that the vacuum cleaner can end up in has to be equal to 1. For example if the vacuum cleaner is next to a sofa it can decide to either avoid it or try to get under the sofa and clean the area there. If it decides to go under the sofa there is 20% that if will not fit there and get damaged, 70% probability that it will clean it successfully and 10% probability that it will clean the area but get stuck there. So if we sum probabilities of all possible states after taking particular action it has to be 1. In our case 0.2 (next state is damaged) + 0.7 (next state next to a sofa after cleaning the area under it) + 0.1 (are is cleaned but it is stuck under the sofa) is equal to 1.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain it is. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&amp;lt;math&amp;gt;\pi (a | s) = P[A_t = a | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Policy is a function for mapping from each states s from all possible states to an action a from all possible actions. In other words it optimizes the best action that should be taken from each state that an agent can end up in, in order to maximize the return function mentioned above (to get as big as reward as possible when facing possible infinite horizon). Because ending in some state is often not deterministic we follow a random path when we follow the policy. When defining a policy i need to use expected utility because of the randomness.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(s) = \sum_s' T(s, a, s') [R(s, a, s') + \gamma V_\pi(s')]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This function tells me the value of policy &amp;lt;math&amp;gt;\pi&amp;lt;/math&amp;gt; applied at state s, which is equal to sum over all possible states s' i can end up in where we sum expected immediate reward R after taking the action a and value of the same policy at state s' after taking discount factor into consideration.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example 1 ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
&lt;br /&gt;
State: at the start of a new round / out of the game&lt;br /&gt;
&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
&lt;br /&gt;
Reward: 4 dollars / 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability. However there is also 33% probability of getting 4 dollars, but after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;br /&gt;
&lt;br /&gt;
Let's now set some policy and calculate value of that policy. So for example our policy is to take action 'stay', when in state 'at the start of a new round'.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} (4 + V_\pi(end)) + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
we know that the expected reward of policy 'stay' is 0, as after the end of the game we won't get any more money, so let's write it into our formula.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = \frac{1}{3} 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 4 + \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;4 = \frac{2}{3} (4 + V_\pi(stay))&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;V_\pi(stay) = 12&amp;lt;/math&amp;gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20233</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20233"/>
		<updated>2021-01-05T22:11:50Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain it is. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&amp;lt;math&amp;gt;\pi (a | s) = P[A_t = a | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Formally it is a probability distribution over actions a available at the current state s.&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example 1 ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
State: at the start of a new round / out of the game&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
Reward: 4 dollars / 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability get to another round where it will get an option to obtain another 4 or 10 dollars. However there is also 33% probability that after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20232</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20232"/>
		<updated>2021-01-05T22:05:09Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
'''Reward''' utility in some units (money, points etc.) after taking an action&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain it is. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&amp;lt;math&amp;gt;\pi (a | s) = P[A_t = a | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Formally it is a probability distribution over actions a available at the current state s.&lt;br /&gt;
&lt;br /&gt;
== MDP Applications ==&lt;br /&gt;
&lt;br /&gt;
=== Example 1 ===&lt;br /&gt;
Imagine we play a game. The game is played in rules. In each rule you get to decide whether you want to quit or continue. If you quit you get 10 dollars and the game is over however if you continue i will give you 4 dollars and roll a 6-sided dice, if the dice results in 1 or 2 the game is over otherwise we start another round.&lt;br /&gt;
First let us list all MDP variables according to terminology.&lt;br /&gt;
&lt;br /&gt;
Agent: a robot&lt;br /&gt;
Enviroment: in this case it is not important for the process&lt;br /&gt;
State: at the start of a new round&lt;br /&gt;
Actions: to quit, to continue&lt;br /&gt;
Reward: 4 dollars / 10 dollars&lt;br /&gt;
&lt;br /&gt;
When a robot is in state when a new round starts, it has 2 actions it can takes, either it can quit the game, or continue. If the robot quits the game it is 100% certain that it will get 10 dollars. If the robot decides to take risk it will get 4 dollars and with 67% probability get to another round where it will get an option to obtain another 4 or 10 dollars. However there is also 33% probability that after rolling a dice, the dice will result in 1 or 2 a thus quits the game.&lt;br /&gt;
&lt;br /&gt;
In this example we can see Markov property in a way that current state depends solely on previous state. So for us it is not important in what nubmer the dice resulted 2 rounds ago, the most important for us is that in last round dice resulted in number 4 and it gave us green for next round.&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20231</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20231"/>
		<updated>2021-01-05T21:16:56Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable. It is categorized as dicrete-time stochastic process. Markov decision process is named after Russian mathematician Andrey Markov and is used for optimization problems solved via dynamic programming.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
The difference between Markov Process and Markov Chain is that Markov Process is in fact an extension of Markov Chain. Markov Process is enriched with rewards (giving motivation) and actions (providing agent with choices). So if reward is zero and only one action exists for each state, Markov Process is reduced to Markov Chain.&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Without actions state transitions would be stochastic (random). Now the agent can control its own fate to some extent.&lt;br /&gt;
&lt;br /&gt;
=== Return ===&lt;br /&gt;
&amp;lt;math&amp;gt;G_t = R_{\text{t+1}} + \gamma R_{\text{t+2}} + ... = \sum_{k = 0}^\infty\gamma^k R_{\text{t+k+1}}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We establish sum over all returns in order to avoid situation when we pick an action that will give us decent reward now but will miss better strategy in long term. In the formula we use discounting factor because futher the reward is less certain it is. Value of discounting factor is with the interval [0, 1]. Further into future we go closer to 0 the discounting factor is.&lt;br /&gt;
&lt;br /&gt;
=== Policy ===&lt;br /&gt;
&amp;lt;math&amp;gt;\pi (a | s) = P[A_t = a | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Policy is the thought behind making some decision. Formally it is a probability distribution over actions a available at the current state s.&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20230</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20230"/>
		<updated>2021-01-05T20:36:49Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_{\text{t+1}} = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_{\text{t+1}} | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by &amp;lt;math&amp;gt;(S, P, R, \gamma)&amp;lt;/math&amp;gt;, where S are states, P is state-transition probability, R_s is reward and &amp;lt;math&amp;gt;\gamma&amp;lt;/math&amp;gt; is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;br /&gt;
&lt;br /&gt;
=== Markov Decision Process ===&lt;br /&gt;
We get Markov Decision Process basically by adding action variable A into Markov Process and Markov Reward Process.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;P_{\text{ss'}}^a = P[S_{\text{t+1}} = s' | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;R_s^a = E[R_{\text{t+1}} | S_t = s, A = a]&amp;lt;/math&amp;gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20181</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20181"/>
		<updated>2021-01-01T23:20:10Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_t+1 = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Markov Reward Process ===&lt;br /&gt;
&amp;lt;math&amp;gt;R_s = E[R_t+1 | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Markov Reward process is defined by '''(S, P, R, \gama)''', where S are states, P is state-transition probability, R_s is reward and \gama is the discount factor. R_s is the expected reward of all possible states that one can transition from state s. By convention, it is said that the reward is received after the agent leaves the state and hence, regarded as R_t+1. For example, let's consider already mentioned automatic vacuum cleaner, if it decides to go cleaning near to a wall there is 10% probability that it will crash what will bring utility of -5, however if it manages to avoid the wall and will clean the area sucessfully (probability is 90% as there are only two probable outcomes) it will get the utility 10. Hence the expected reward that the vacuum cleaner gets when going close to the wall is '''0.1 * (-5) + 0.9 * 10 = (-0.5) + 9 = 8.5'''&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20180</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20180"/>
		<updated>2021-01-01T22:09:35Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_t+1 = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider an example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20179</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20179"/>
		<updated>2021-01-01T22:08:59Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_t+1 = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider the example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20178</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20178"/>
		<updated>2021-01-01T22:08:24Z</updated>

		<summary type="html">&lt;p&gt;Sára: /* Markov Process/Markov Chain */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_t+1 = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;br /&gt;
A Markov Process is defined by (P, S), where S are the states and P is state-transition probability. It consists of a sequence of random states s1, s2 ..., where all states obey Markov Property. Pss' is the probabilty of jumping to state s' from current state s.&lt;br /&gt;
Let's consider the example of an automatic vacuum cleaner. When it is next to a wall there is probability of 10% that it will crash it and 90% probabilty that it will change direction and proceed with cleaning. So the probability of state s' (crashed in our case) is 0.1 with respect to current state (next to a wall).&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20177</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20177"/>
		<updated>2021-01-01T22:00:36Z</updated>

		<summary type="html">&lt;p&gt;Sára: /* Markov Process/Markov Chain */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_t+1 = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20176</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20176"/>
		<updated>2021-01-01T21:59:29Z</updated>

		<summary type="html">&lt;p&gt;Sára: /* Markov Process/Markov Chain */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_(t+1) = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20175</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20175"/>
		<updated>2021-01-01T21:58:09Z</updated>

		<summary type="html">&lt;p&gt;Sára: /* Markov Process/Markov Chain */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_t_+_1 = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20174</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20174"/>
		<updated>2021-01-01T21:57:34Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move around the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;br /&gt;
Markov property says that current state of the agent (for example a Robot) depends solely on the previous state and doesn't depend in any way on states the agent was in prior the previous state.&lt;br /&gt;
&lt;br /&gt;
===Markov Process/Markov Chain===&lt;br /&gt;
&amp;lt;math&amp;gt;Pss' = P[S_t+1 = s' | S_t = s]&amp;lt;/math&amp;gt;&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20162</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20162"/>
		<updated>2020-12-28T09:18:50Z</updated>

		<summary type="html">&lt;p&gt;Sára: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move arounf the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;br /&gt;
&lt;br /&gt;
== Characteristics ==&lt;br /&gt;
&lt;br /&gt;
===Markov Property===&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
	<entry>
		<id>http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20160</id>
		<title>Markov decision process</title>
		<link rel="alternate" type="text/html" href="http://www.simulace.info/index.php?title=Markov_decision_process&amp;diff=20160"/>
		<updated>2020-12-28T07:42:19Z</updated>

		<summary type="html">&lt;p&gt;Sára: Created page with &amp;quot; == Introduction == Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable....&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
== Introduction ==&lt;br /&gt;
Markov decision process is a mathematical framework used for modeling decision-making problems when the outcomes are partly random and partly controllable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Terminology===&lt;br /&gt;
'''Agent:''' an agent is the entity which we are training to make correct decisions (we teach a robot how to move arounf the house without crashing).&lt;br /&gt;
&lt;br /&gt;
'''Enviroment:''' is the sorrounding with which the agent interacts (a house), the agent cannot manipulate its sorroundings, it cannot only control its own actions (a robot cannot move a table in the house, it can walk around it in order to avoid crashing).&lt;br /&gt;
&lt;br /&gt;
'''State:''' the state defines the current situation of the agent (the robot can be in particular room of the house, or in a particular posture, states depend on a point of view).&lt;br /&gt;
&lt;br /&gt;
'''Action:''' the choice that the agent makes at the current step (move left, right, stand up, bend over etc.). We know all possible options for actions in advance.&lt;/div&gt;</summary>
		<author><name>Sára</name></author>
		
	</entry>
</feed>